Sora:AI 视频革命与社交体验

Sora: The AI Video Revolution and Social Experience

比尔·皮布尔斯 Bill Peebles · Unsupervised Learning · 2025-11-03 · 约 63 分钟 · 原视频 ↗

打开互动全文版(中英对照 + 朗读 + 问答)→

本期速览 · Overview

OpenAI 的 Sora 团队讨论意外的病毒式成功、将其打造成社交体验的产品决策,以及包括新颖物理发现在内的未来里程碑。

OpenAI's Sora team discusses the unexpected viral success, product decisions behind making it a social experience, and future milestones including novel physics discoveries.

要点 · TL;DR

核心观点 · Key points

反共识 · Contrarian takes

本期章节 · Chapters(共 37)

全文 · Full transcript(中英对照)

引言与Sora发布 Introduction and Sora's Launch

Host

我觉得我请来了 AI 生态圈的主角们。人们非常有创造力。这是一个社交网络。所以,社交网络是关于人和关系的。即使你把创作者和消费者的比例改变 1%,那也会对世界产生巨大影响。名人和版权所有者正在快速进化,理解这项技术以及如何使用它。Jack 在发爆款视频。顺便说一句,我认为第一个通过视频模型模拟某种现象而带来的科学突破将是一个疯狂的里程碑,对吧?你对那可能发生的时间有预测。如果到那时我们还没有取得这样的突破,我会很震惊。Sora 在过去一个月风靡全球,它在应用商店排名第一,催生了无数搞笑视频。今天我有幸与 Bill、Rohan 和 Thomas 交谈,他们是 OpenAI 的 Sora 背后三位核心人物。这是一次非常有趣的对话。我问了他们所有关于 Sora 引发的问题。我们讨论了他们对 Ben Thompson 关于 Sora 看法的反应,以及影响他们将其打造成社交体验的关键产品决策。我们讨论了视频模型的现状,过去几年的进展,以及未来的里程碑,包括新颖的物理学发现及其时间线预测。我们还谈到了视频模型如何融入更广泛的 OpenAI 生态系统,包括未来的 ChatGPT。与这三位才华横溢的人谈论视频生态系统的所有事情非常有趣。我想大家会非常喜欢。话不多说,有请 Sora 团队。

I feel like I have the main characters of the AI ecosystem. People are very creative. It's a social network. So, social networks are about people and relationships. Even if you change the ratio of creator to consumer by 1%, like that is a massive impact on the world. Celebrities and rights holders are going through like a very fast evolution of understanding this technology and like how to use it. Jack is posting bangers. By the way, I think the first scientific breakthrough that comes as a result of simulating some phenomenon in video models would be an insane milestone, right? You have a prediction of when that might be. I would be shocked if we are not making breakthroughs like this by Sora has taken the world by storm this past month. It's number one in the app store has spawned endless hilarious videos and I had the privilege today of speaking with Bill Rohan and Thomas three of the main folks at Sora behind this magic at OpenAI. It was just a really fun conversation. I got to ask them everything that I've been thinking about that Sora spawned. We talked about their reactions to Ben Thompson's take on Sora as well as kind of the key product decisions that factored into how they thought about making it into a social experience. We talked about the state of video models, how they progressed over the past years and what's next future milestones including novel physics discoveries and their predictions for timelines on those. We also hit on how video models in general will fit into the broader OpenAI ecosystem including chat GBT over time. Just a ton of fun to get to talk with these three brilliant folks about all these things in the video ecosystem. I think folks will really enjoy it. Without further ado, here's the Sor team.

Host

各位,我非常兴奋能进行这次对话。非常感谢你们来到这里。

Guys, I'm so excited to do this. Thank you so much for coming in here.

Bill Peebles

谢谢邀请我们。

Thanks for having us.

Host

我觉得我请来了 AI 生态圈的主角们,作为播客主持人,这是一件很有趣的事情。所以,谢谢你们。我觉得我们今天有很多不同的话题要聊。也许从开始说起,我听过很多关于 ChatGPT 发布的故事,它原本没被期望成为什么大事件,却突然变成了一个被广泛使用的消费产品。我很好奇 Sora 是否也是如此?或者说,你们预料到这种反应了吗?第一天有什么有趣的故事吗?

I feel like I have the main characters of the AI ecosystem, which is like as a podcast host, a fun thing to get. So, thank you guys. I feel like a bunch of different things we'll want to hit on today. You know, maybe to start, I feel like there's been all these stories told of the chat GBT launch and how it was kind of not expected to be anything big and out of nowhere turned into this massively used consumer product. I'm curious whether the same was true of Sora or I guess like did you guys kind of expect this reaction? Any kind of fun stories from that first day?

Bill Peebles

我的意思是,我没想到会在应用商店排名第一长达一个月。但这是你回顾时会觉得研究团队做得非常出色的事情。OpenAI 很擅长制造这种病毒式传播的时刻,很高兴我们做到了,但它超出了我的预期。

I mean I didn't expect to be number one in the app store for like a month. But it's one of those things you look at in retrospect like the research team absolutely cooked. We have a knack at OpenAI of creating these viral moments really well and so happy to have done it but I think it surpassed my expectations.

Host

我觉得 Bill 还有一种无望的乐观。其实不是无望。我记得有个梗,发布几天后我看到它在应用商店排名第三,然后 Bill 就要求拿第一。我想 Sam 把你关进笼子之类直到你拿到第一。

I think Bill also has a hopeless optimism about him. It's not hopeless. Actually, I think it was like there was a meme. I saw it the first it was a couple days after launch where it was like number three in the app store and then Bill was like demanding number one. I think Sam was putting you in the cage or something until you were number one.

Bill Peebles

所以我最像 2.5K 并把它变成了现实。

So my most like 2.5K and willed it to life.

Host

你知道,我一直在很多不同的社交媒体产品上工作,在 Instagram 待过一段时间。是的。在那个生态系统中,有些成功,有些失败。所以,我倾向于非常务实和现实。但我必须说,那种乐观实际上非常重要。我们最终在应用商店排名第一,这完全超出了我的预期。但看到这一切非常令人高兴。

You know what I'm always I've worked a lot of different social media products in Instagram for a while. Yeah. and some successful, some failures within that ecosystem. So, I tend to be very grounded and realistic and all this sort of stuff. But I have to say like that optimism is actually very important. We were we ended up being number one on the app store which was wildly crazy for my expectations. But it's delightful to see.

Host

所以,我想这方面你们需要预留一些 GPU 算力,对吧,基于你们的预期。我一直想知道你们到底是怎么算出来的?

So, I imagine there's an aspect of this where you kind of have to reserve some sort of GPU capacity, right, to based on what you expect. And so, I always wonder like how in the world do you figure that out?

Bill Peebles

你知道,这其实不是一门科学。我认为整个行业的人都感受到算力紧张,不仅仅是 OpenAI。当然,当我们推出一个全新的产品界面,而且是像视频这样计算密集型的,就需要公司其他部门做出一些牺牲。我认为 OpenAI 的一个优点,也是 Sam 的功劳,就是每个人都觉得彼此的成功息息相关。当然,当 ChatGPT 发布惊人的新视觉生成功能时,我们希望确保它们很棒,人们在聊天中获得出色体验;同样,公司也真心希望 Sora 成功,人们明白我们需要进行这些新尝试,这对公司的长期发展很重要——不仅要有一个非常成功的面向消费者的语言模型产品,还要在视频领域取胜,每个人都愿意在战时贡献自己的一份力量。

You know, it's not really a science. It's a lot of, I think everyone feels compute across the entire industry, not just with OpenAI. And certainly when we launch an entirely new product surface, right, that is as compute-intensive as video, it requires some pain elsewhere in the company. I think one thing that's really nice about OpenAI and like to Sam's credit, everyone kind of feels invested in everybody's success across the board, right? Certainly when ChatGPT releases like amazing new visual generation features we want to make sure those are great and people are having an awesome experience in chat and likewise the company really feels invested in Sora's success people see that we need to take these new bets and it's important for the company's longevity not just to have a super successful LM facing consumer product but also the winning one in video and everyone's willing to chip in do your part in wartime.

Sora独立应用 Sora as an Independent App

Host

将 Sora 作为独立应用的想法是什么时候产生的?从一开始就很明显吗?

When did the idea to create Sora as an independent app come about? Was that always obvious from the beginning?

Bill Peebles

这并不总是显而易见的。我认为有一个清晰的愿景,并且对于这项技术能做什么以及如何改变用户生成内容和媒体整体,一直有很高的雄心。如果你把它与我们实际使用模型时的体验结合起来——比如在多人模式中有一些非常特别的东西,我们在 Sora 2 之前就看到了这种涌现行为,即 AI 使人们能够极快地参与趋势并创造趋势,就像我们在其他视频生成软件中看到的那样,这越来越多地涌入社交媒体。然后,当我们推出图像生成功能时,它在 ChatGPT 中爆火,人们把自己置身于这些惊人的场景中。所以,看到自己出现在这些视频中,我们想更进一步,比如和朋友们一起做。我们内部感受到了魔力,一旦我们感受到了,这就是一个非常容易的决定。

It wasn't always obvious. I think there was a clear vision and there has been a great level of ambition in terms of what this technology can be and how it can transform UGC and media generally. And so then if you pair that with what we were actually experiencing with the model with like hey there's something very special in multiplayer mode we're seeing this emergent behavior even before Sora 2 in terms of what is AI enabling people can participate in trends extremely quickly and create these trends extremely quickly like we've seen that with other video generation software in terms of this is flooding social media more and more. And then with image gen when we launched image gen it's amazing it blew up in ChatGPT but people were just putting themselves in these amazing scenes and so something about seeing yourself casting yourself in these videos and we just wanted to take that one step further like do it with your friends and we felt magic internally and I think once we felt that it was a very easy decision.

Host

还有一点要补充的是,这些界面未来可能会更接近,这并非不可想象,但 ChatGPT 目前感觉非常像单人游戏,这在某种程度上是神圣的。如果你不知道在 ChatGPT 中做的事情是公开的还是私密的,可能会有点不协调。所以我们也对此很敏感。是的,这实际上是一个漫长的旅程。

The one thing to add also is it's not like inconceivable that maybe these surfaces come closer together, but ChatGPT does feel very single player and that's sacred in a way right now like it can feel a little jarring if you don't know when you do something in ChatGPT is this public or not. So we also we're sensitive to that as well. Yeah, there's actually a long journey.

最古怪原型:ChatGPT内社交流 The wackiest prototype: a social media stream inside ChatGPT

Bill Peebles

我非常高兴它最终用在了 Sora 上,但对我来说并不明确。产品之旅总是非常……

I'm very delighted it ended up in Sora, but it was not clear to me. Product journey is always extremely...

Host

你扔掉的最奇怪的东西是什么?

What was the strangest thing you threw out?

Bill Peebles

我们有一个原型,就是在 ChatGPT 里做一个纯粹的社交媒体信息流,内部试用。我们看到的一个现象是:想到社交,第一反应就是,我们有 ChatGPT,为什么不加一些社交功能看看效果?内部开始用的时候,正好是在图像生成之前。我们有那种 Reddit 风格的线程式链条。我们会看到很长的链条,有人放一张图,然后另一个人说,“现在它抽雪茄了。”比如是一只鸭子。一只抽雪茄的鸭子。现在它倒过来了。现在它变红了。看着它演变真的很酷。就像这些小 meme 链条。那看起来非常特别。想想 Rohan 说的,这其实是如果没有生成式 AI 就做不到的事情,因为创作太难了,混搭、即兴发挥都太难了。所以那是我们做过的最古怪的东西。

We had a prototype of just a pure social media stream inside of ChatGPT, and we're trying it out internally. One of the things we saw: think about social, first thing that comes to mind, we have ChatGPT, why don't we just put some social features in it, see what it looks like? When we started using it internally, it happened to be right before the image generation. We had these threaded, Reddit-style chains. We would see very long chains where somebody would put up an image, and then somebody would be like, "Okay, it's smoking a cigar now." Say it's a duck. It's a duck with a cigar. Now it's upside down. Now it's red. It was really cool to see it evolve. It was like these little meme chains. That seemed like something very special. Thinking about what Rohan was saying, that's actually something you can't really do without generative AI because it's too hard to create, too hard to remix, to riff on something. So that was the wackiest one we had.

杀手功能:客串自然涌现 The killer feature: cameos emerge organically

Bill Peebles

经过漫长曲折的产品之旅,Sora 模型变得相当不错。我们一直想用它们做些真正有野心的事情,大规模部署。幸运的是,它就像手套一样合适。图像生成很酷,但视频在某些方面更酷,尤其是在这个背景下。

Through this long, snaky product journey, the Sora models were getting pretty good. We always wanted to do something really ambitious with them, to really mass deploy them. Fortunately, it just kind of fit like a glove. Image generation is very cool, but videos are even cooler in some ways, especially in this context.

Host

听起来从图像工作开始,你们就隐约知道客串会是这些强大模型的杀手锏功能。

It sounds like from the image work, you kind of knew cameos would be this killer feature for these powerful models.

Bill Peebles

我不知道我们是否知道。这真的很有趣。事后看来,这是显而易见的事情,但我觉得这个想法在研究团队构建模型时就已存在。直到我们真正有了它……我们团队的一位工程师 Bobo 说:“每个人都给我发一段视频,说‘嘿 Sora,我是 Phil’。”对。“嘿 Sora,让我活过来。”我们就在 Slack 线程里做了。他在后端上传了所有这些。然后你就可以直接标记那个人。我们之前讨论过这个想法。然后终于,那天一切都变了,我想。我们都开始互相客串。当时我们还没有一个动词来形容这个。我们只是开始互相标记。它在我们注意到之前就发生了。不是那种“砰,就是它了”的感觉。就在几天内,我觉得我们都在说:“我们的信息流全是客串。怎么回事?”我们都在用这个应用。

I don't know if we knew. It's really interesting. It's one of those things in retrospect. Obviously, but I think this idea was in the ether from when the research team was actually building the model. It wasn't until we actually had it... One of our engineers on the team, Bobo, was like, "Everyone send me a video of you saying, 'Hey Sora, it's Phil.'" Yeah. "Hey Sora, bring me to life." We did it in a Slack thread, literally. He uploaded all these in the back end. Then you could just tag the person. We had talked about this idea. And then finally, everything changed that day, I think. We all just started cameoing each other. We didn't have a verb for this at the time. We just started tagging each other. It happened before we even noticed. It wasn't like "boom, this is it." Just in a couple days, I feel like we were like, "Our feed is just cameos. What's going on?" And we were all using the app.

Host

我们确实有过一个时刻,心想:“这不好吗?”但等等。我们之前有有趣的内容,现在全是人……等等,不,这其实很棒。这正是我们需要的。

We actually had a moment where we're like, "Is this bad?" But yeah, wait a second. Did we just... we had interesting stuff before, and now it's just people... like, wait a second, no, that's actually amazing. That's what we need.

Bill Peebles

是的。

Yeah.

用户行为与涌现能力 Surprising user behaviors and emergent capabilities

Host

构建这些产品最酷的一点是,你发布第一天,人们就会立刻以你从未预料到的无数种方式使用它。你见过的最令人惊讶的使用方式是什么,或者你之前不知道存在的功能?

One of the coolest parts of building these products is you release it day one, and people start immediately using it in a million different ways that you never intended. What's the most surprising way you've seen people use it, or capabilities that you didn't know actually existed?

Bill Peebles

有些人把自己放在励志场景中,想象自己处于想要的情境。我觉得每当我看到风格转换……这个词可能太小了,但你拿一个东西,拿一个人。有一个评论我本来不相信会成功,结果却成了被混搭最多的视频之一:一个孩子打开圣诞礼物,里面是一个 Bill Peebles 动作人偶。看起来像 Bill Peebles。这就是让我留在排行榜视频上的原因。我大概是第 25 名左右。Bill Peebles 动作人偶。模型能理解这个,理解上下文,仅凭几个提示就能把你放在一个完全不同寻常的场景中,这很了不起。我很惊讶模型能涌现出这种能力。我在信息流里经常看到这类东西。不完全一样,但总是围绕这个主题的变体。非常酷。还有那些大规模转换的……声明类的?我喜欢电子游戏类的。我觉得现在这方面还有点未被充分探索,比如一个 LucasArts 冒险游戏,但主角实际上是 Rohan 或 Bill 之类的。那些我简直爱不释手。立刻就想看看那些视频里的秘密表情符号是什么。

Some people were putting themselves in motivational situations, envisioning themselves in situations they wanted to be in. I think anytime I see a style transform... it's maybe too small a word for this, but you take something, take a person. There's one review that I couldn't believe would work and ended up being one of the most remixed videos: a kid opening a package for Christmas, and it's like a Bill Peebles action figure. It looks like Bill Peebles. This is what's keeping me on the leaderboard video. I'm like 25th or something. Bill Peebles action figure. It's remarkable that the model understands that, and understands the context from just a couple of prompts, can put you in a completely unusual scene. I was amazed that emerged from the model. I see those types of things in my feed all the time. Not exactly the same, but it's always some riff on that. That's very cool. And anything that's massively transformed... the claimation ones? I love the video game ones. I think it's a little underexplored right now, where you have like a LucasArts adventure, but it's actually Rohan or Bill or something. Those I just eat up. Immediate like, see what the secret emoji is on those ones.

Host

有些人简直疯了,比如 Ka Kaumu Matsumara,就像……我们甚至都没接近那种风格输出水平。所以,看到世界的创造力真是太好了。

Some people are just getting insane, like the Ka Kaumu Matsumara, just like... things we didn't even get close to that level of stylistic output. So, just seeing the creativity of the world.

Bill Peebles

是的。

Yeah.

Host

是的。我认为在我们推出 Storyboard 之后,它可以让你生成最长 25 秒的视频,那时我的质量门槛真的提高了。我实际上很惊讶,用这个模型一次生成,就能得到难以置信的连贯故事。这是你用 Sora 1 尝试 100 次都得不到的。这真的是 Sora 2 的新东西。它反映了智能的阶跃式提升。

Yeah. I think after we launched Storyboard, which lets you generate up to 25 seconds of video, that's when the quality bar for me really spiked. I'm actually surprised how in one shot out of this model, you can just get unbelievably cohesive stories. This is something that you cannot even get with like 100 attempts at Sora 1. This is really a new thing in Sora 2. It just really is reflective of the step function increase in intelligence.

Sora重创作轻消费 Sora's focus on creation vs. consumption

Host

在 TPN 上,我记得你分享过大约 70% 的 Sora 用户是创作者,这令人难以置信。我肯定这个比例仍然很高。关于 Sora 的讨论中,我觉得有趣的一点是,当它最初发布时,Ben Thompson 写了一篇文章,他说:“我对 Sora 持怀疑态度。大多数人不想创作,他们想消费。”这是其他所有产品的行为模式。然后我觉得他给我打了个电话。他说:“好吧,我错了。这是个很酷的产品。”我很好奇你的反应是什么,以及你认为这种创作水平是否会持续下去。

On TPN, I think you shared that like 70% of Sora users are creators, which is mindboggling. I'm sure it's still super high. One thing I thought was interesting around the discourse around Sora is when it was initially released, I think Ben Thompson wrote this piece and he was like, "I'm skeptical of Sora. Most people don't want to create, they want to consume." That's been the behavior of every other product. And then I think he issued me a call. He was like, "Okay, I was wrong. This is a cool product." Curious what your reaction was to that, and you think this level of creation kind of persists.

Bill Peebles

是的,我的意思是,我们确实从头开始设计这个应用,专注于创作方面,这是我们进入这个领域的核心假设。现有的社交媒体网络有很多很棒的地方,人们从中获得很多快乐,这可以来自消费端。但同时,我认为最严重的伤害和最坏的情况来自无休止的刷屏。我们真的想弄清楚如何避免这种非常反乌托邦的未来:一个完全针对你想看的内容优化的垃圾信息流,与朋友脱节,感觉非常孤独和孤立。

Yeah, I mean we really designed the app from the ground up to be focused on the creation side, and this was really our core hypothesis going into it. There's so many ways existing social media networks are really amazing and people get so much joy out of them, and that can come from the consumption side. At the same time, I do think probably the worst harms and the worst case scenarios come from doom scrolling. We really wanted to figure out how to avoid this very dystopian future of just a slop feed that's optimized exactly to what you want to watch, very detached from your friends, feels very lonely and isolating.

客串与推荐系统设计 Cameo and Recommender System Design

Bill Peebles

我们在这方面做了全面改进。我认为最关键的是 Cameo 功能。Cameo 真正让生成内容对你来说变得个性化,并且非常人性化,这是文本转视频或那些简单的提示方式可能做不到的。此外,Thomas 在推荐系统方面做了大量工作。如果你不小心或者不是善意的参与者,这些系统很容易失控——如果你把推荐系统设计成最大化点击诱饵和参与度,就会助长学生刷屏的行为。Thomas 做了很多开创性的工作,重新思考了现代推荐系统栈在这个产品中应该是什么样子,使其更偏向创造力而非消费。

And we kind of tackled this across the board. So like really I think the single most important thing we did there was really with Cameo. Like Cameo is really what makes a gen feel personal to you and makes it feel very human in a way that just kind of like text to video or just like all these very simple ways of prompting these models maybe doesn't. Um and then we also did a lot of work Thomas specifically on the recommender side. Right. So like really the way that these things really go off the rails if you're not careful or being kind of like a good faith actor is if you design this rec sys in a way which is sort of like maximally clickbaity/engaging and like kind of incentivizes the student scrolling behavior and Thomas did a lot of like very pioneering work to really rethink what like a modern rec sys stack looks like in this product to oriented towards creativity over consumption.

Host

是的。这实际上在某种程度上是自我实现的,或者说是一个循环。但你可以参与网络,这意味着你也可以鼓励——这现在成了一个可以优化的目标。我认为这是一个非常健康的目标:如果你真的决定“我要混搭这个生成内容”,那是一种非常主动的行为,让你进入一种创造模式,而不是纯粹的消费模式。这就是推荐系统背后的理念:因为创作如此容易,我们实际上可以鼓励人们去创造,这是通常不会优化的方向。我在 Instagram 工作过,当时你看到一张照片或视频,然后打开相机拍照,再把它归因回原始照片,这几乎是不可能的。你只能得到一点点信号。但在这个世界里,你可以从一切中获得信号。所以,我认为这是一种非常微妙但有趣的行为。

Yeah. And some of that is actually it's kind of self-fulfilling in a way or like a circle. But the fact that you can participate in the network means that you can also encourage like you that's now an optimization objective that can be in there. And I think it's a very healthy one if you're actually making a decision to oh I'm going to remix this generation. Um that's like a very active behavior. It puts you in a very creative mode that's not that whole just consumption-oriented thing. And so that was the idea behind the rec sys is like because it's so easy to create, we can actually encourage people to create in a way that you don't normally optimize for because uh I mean having worked at Instagram, the idea that you saw this photo or this video and then you help you know opened your camera and then you took a photo and you have to attribute it back to the original photo you saw. It's impossible. Or it's very close to impossible. You can get a little bit of signal out of that. Uh but in this world you can get signal out of everything. So, I think it's a very nuanced but kind of interesting sort of behavior that you can see.

Ben Thompson反馈弧 Ben Thompson's Feedback Arc

Bill Peebles

是的。关于 Ben Thompson,我觉得他的反馈历程很有意思。他第一次给出反馈时,我真的很受打击。但他的历程是合理的,对吧?他是一名消费者。冷启动网络是一个非常困难的问题,所以信息流还没有完全调好。他也是早期用户,信息流还没有经过足够的时间沉淀,所以没那么有趣。作为一个纯粹的消费者而不是创作者,他一开始就跳进去了。我记得他有一句评论,说它就像 Sam Altman 的广告之类的,公平地说,信息流确实有一阵子给人那种感觉。我觉得这挺有趣的。但后来当他回到他的播客时,他提出了一个很好的观点:即使你把创作者与消费者的比例改变 1%,也会对世界产生巨大影响,对吧?那会是一个完全不同的产品。我认为这很好地总结了为什么我们认为 Sora 很特别,以及他第一次可能错过了什么,因为他日常生活中是一个纯粹的消费者,而不是创作者。

Yeah. And like on Ben Thompson, I think his arc, you know, I was crushed when he first gave his first round of feedback. But his arc makes sense, right? Like he is a consumer. Cold start network is very hard problem. So the feed wasn't super dialed in. He's an early adopter too. So early adopter without like a feed that has had any time to bake. It wasn't that interesting. So, as a pure consumer and not a creator, I think he jumped in there. I think he had a quote about it just being like a Sam Altman advertisement or something, which the feed did feel like that for a little bit to be fair. I mean, I thought it was funny, but um but then I think when he had his kind of like, you know, he brought it back to his mayulpa, he had a good point which is like even if you change the ratio of creator to consumer by 1%, like that is a massive impact on the world, right? And that's like such a different product. Um, and I think that was like a good summary of why we think Sora is special and what he might have missed the first time around because he is particularly like a pure consumer and not a creator in his everyday life.

人类创作vs机器人创作 Human Creation vs Bot Creation

Bill Peebles

我确实认为,我们在 Sora 的一些早期原型中也探索过这一点。但人类创作和机器人创作之间有一个非常根本的区别。这很难体会,但如果你把今天的 Sora 信息流去掉 Cameo,再隐去发布者,我认为它会非常无趣。关键的一部分是,有人看过这个内容,并决定给它盖上自己的认可印章。他们也参与了创作过程。这很容易理解。我见过一些完全由 AI 驱动的社交网络,但它们的吸引力只能维持五分钟。我收到一个机器人的点赞,我会想:这点赞是什么意思?我不太理解。但如果 Ron 喜欢我的内容,哦,是的,我知道那意味着什么——嘿,你看到了我的东西,你喜欢它吗?对吧?人类创作的东西有一种非常直观的感觉,因为他们投入其中。

I do think it's also we play with this in some of our early prototypes off Sora. But the idea that there's a human creating it versus there's like a bot creating it is a very very foundational difference. It's like kind of hard to appreciate, but if you took the Sora feed today, maybe excluding cameos and you just stripped the person that posted it, it would be very uninteresting in my opinion. Like part of the essential thing is that somebody has looked at this and decided this is like kind of like they're putting their stamp of approval on it. They're also involved in the creative process. And I think that is like easy to think about. I've seen a few of these social networks that are completely AI-driven, but they last for like five minutes of interest. I get a like from a bot and I'm like, what does it like mean? I don't really understand. Whereas like if Ron likes my content like oh yeah you know I know what that means like hey you saw my thing like did you like it like so right something almost that's so visceral about the fact that a human created it because they're in the they're in it right other other tools.

名人与权利方演变 Celebrities and Rights Holders' Evolution

Host

随着 Sora 的推出,还有一件非常有趣的事情是,我觉得名人和版权方正在快速进化,理解这项技术以及如何使用它。你能谈谈过去一个月他们的历程是怎样的,以及大多数人现在处于什么阶段吗?

One thing that's also been really interesting to see as Sora has been out is just like I feel like celebrities and rights holders are going through like a very fast evolution of understanding this technology and how to use it. Can you talk a little bit about like what that journey's been like over the past month and where most folks are today?

Bill Peebles

是的,我们自推出以来与社会的各个方面都进行了交流。如果你把时间倒回一个月前,世界上大多数人真的不知道视频生成是什么。当然他们也不知道 Sora。随着时间的推移,我们有机会与这些人坐下来,倾听他们对这个平台兴奋的地方。这里有很大的价值主张,尤其是对版权方。我们昨天推出了角色 Cameo,对吧?你可以想象,如果你有一些人们喜爱的 IP,现在任何孩子都可以生成一个包含该角色的生成内容。这对版权方来说将是巨大的。同时我们也听取了他们的担忧,比如确保他们对角色的出现方式有大量发言权。他们想要限制,不希望它变成完全的自由放任。

Yeah, I mean we've chatted with kind of every facet of society since launching this and you know I think if you rewind the clock like a month ago you know most people in the world really did not know video gen was a thing. Certainly they didn't know Sora. Um uh over time, you know, we've really like just had the opportunity to sit down with these people and hear out both where they're really excited about this platform, right? I mean, there's a huge value proposition here, especially for rights holders. We launched character cameos yesterday, right? You can imagine if you have some IP that people love and any kid now can generate a gen with them in that character. I mean, that's going to be huge for these rights holders. Uh and we're also hearing them out on their concerns, right? Like making sure that they have a lot of input on how their characters show up. They want restrictions. They don't want it to just kind of be a free-for-all.

Host

是的,我认为这是一个相当巧妙的 UI,用来设置基本规则。我确信真正的名人有更复杂的流程,如果有名人特别担心加入平台,我们会一步步引导他们如何设置这些东西。但当我们向人们介绍我们计划中的内容以及已经集成的功能,比如角色 Cameo、Cameo 限制等,实际上有很多兴奋点。而且,就在今天早些时候,我们宣布开始为 Sora 引入变现功能。我们设想未来这是一个很好的宣布方式,比如在你上线前 30 分钟发布消息。这就像在这一集之前保密了好几天。

Yeah, I think a pretty clever UI for setting the basic rails around that. I'm sure the real celebrities have more sophisticated flows and if there's a celebrity who's particularly anxious about getting on the platform, we really step them through exactly like how they should think about setting these things. But when we kind of walk people through what we have on the docket and the features we've already integrated with character cameos, cameo restrictions, etc. There's honestly a lot of excitement around this. And, you know, just earlier today, we announced that we're starting to bring monetization to Sora. We imagine in the future this very nice way to announce that, you know, 30 minutes before you come out, you got to break. This is like holding it for days right before this episode.

变现与早期用户 Monetization and early adopters

Host

我们将开始试点新的方式,让版权所有者通过他们的内容变现,并且会优先考虑那些从一开始就投资平台、现在已经加入的人,我们认为这会带来很棒的结果。

We are going to start piloting new ways for rights holders to monetize their content, and we're going to really prioritize folks who have invested in the platform from the beginning, have gotten on now, and we think amazing things are going to happen from that.

Host

有没有哪些特定的名人或版权所有者让你觉得“哦,他们已经懂了”?我的意思是,马克·库班立刻就把 Cost Plus Drugs.com 加到了他的 Cameo 说明里。我觉得他是第一个想明白这本身可以成为一个很棒的行业的人,从品牌广告之类的角度来看。

Have there been any specific celebrities or rights holders that you're like, "Oh, they already get it." I mean, Mark Cuban immediately added Cost Plus Drugs.com to his Cameo instructions. I feel like he was first to figure out what could be an amazing industry in itself in terms of brand advertisement and stuff like that.

Host

有很多人。我们有一个审核过滤器,阻止你创建公众人物。所以当你走 Cameo 流程时,很多公众人物实际上会被屏蔽,因为我们对冒充行为非常敏感。所以我觉得这很有趣,因为我们收到了很多咨询,他们说“我们没法用你们的产品”。但我们收到的咨询量非常惊人。所以很多人真的在尝试这个,使用它,玩得很开心。沙奎尔·奥尼尔非常有创意。我们和他的团队聊过,他们说“沙克就是玩得很开心,一直在生成内容”。我甚至不知道沙克是否在意他在这个新网络上建立的名声。他只是纯粹在享受创意乐趣。

There have been a lot of people. We have a moderation filter that prevents you from creating a public figure. So when you go through the Cameo flow, a lot of public figures actually get blocked because we are very sensitive to impersonation. So I think it's funny because we get a lot of inbound because they're like, "We can't use your product." But the amount of inbound we got has been amazing. So tons of people are actually trying this out, using it, having an amazing time. Shaq is incredibly creative. We talked to his team and they're just like, "Shaq's just having a blast just generating." I don't even know if Shaq cares about the name he's making for himself on this new network. He is just genuinely having a creative fun time.

Host

顺便说一句,沙克发的都是爆款。

Shaq is posting bangers by the way.

Host

快去关注沙克。

Go follow Shaq.

Host

不过你得打开它才行。

You just got to open it up though.

Host

是的。

Yeah.

与谷歌Meta视频模型竞争 Competition with Google and Meta video models

Host

谷歌和 Meta 都在差不多同一时间发布了类似的产品。大家都在这些视频模型上竞相前进。想知道你怎么看这些发布,以及你认为这个领域会如何发展。显然现在有几家拥有非常好的视频模型。

You've had both Google and Meta release similar kind of products around the same time. Everyone's kind of racing forward on these video models. Wonder what you make of those releases and how you think this space evolves over time. Clearly there are a few folks with very good video models today.

Host

我认为所有这些模型的未来都非常光明。我觉得没有哪个世界能让我们一直保持领先,多亏了比尔的出色工作。但这显然是一种新的媒介和新的创意表达。我认为这个行业会彻底演变,这将成为那个用例的重要组成部分。我认为 Meta、谷歌等公司是对的,我确信他们内部有“红色代码”,或者昨天有人开玩笑说会有“紫色代码”,他们得为这个发明一种新颜色。所以我认为这实际上是正确的,因为他们看到了这里正在出现一些由这种神奇技术创造的东西。不过,到目前为止,我对这些工作的产品化并没有特别深刻的印象。我真的认为这又回到了我们的理念:人与人之间的拥抱。人们非常有创造力。这是一个社交网络。所以社交网络是关于人和关系的。真正推动我们前进的是 Cameo 这个想法,以及将这种创意方面带入你和朋友的生活中。我认为这一点被忽略了。我并不惊讶。我不仅在模型层面不惊讶,而且我完全不惊讶,因为走到这一步是一个非常非线性、曲折的旅程。并不是只有做 Cameo 的想法,还有其他想法,比如在这种语境下“混音”是什么意思?我们有一个一直很奇怪的想法。我觉得没有其他人喜欢这个,但我一直很喜欢,那就是你录一段自己在角落里对 AI 视频做出反应的视频。那是一种反应视频。我们应该把它带回来。

I think the future is very bright for all these models. I think there's no world I think so will stay on top because of Bill's fine work. But this is obviously a new medium and a new creative expression. I think the industry will evolve completely and this is going to be a big part of that use case. I think Meta, Google etc. are right to be like, I'm sure there's a code red or somebody joked yesterday there'd be a code purple or they have to invent a new color for this one. So I think that's actually correct because they're seeing something emerge here that has been created with this magical technology. That said, I haven't been overly impressed with some of the productization of the work so far. And I really think that goes back to our idea of people embracing people. People are very creative. It's a social network. So social networks are about people and relationships. And the thing that really enabled us to move forward is this idea of Cameo and bringing this creative aspect to your life with your friends. And I think that was kind of missed. I'm not surprised. I'm not surprised both on a model level, but I'm not totally surprised because it was a very nonlinear, snaky journey to get to this point. It wasn't like there were ideas of doing Cameo, but there were also other ideas like what does remix mean in this context? We had one that was always weird. I don't think nobody else likes this one, but I always liked it, which was you would take a video of yourself reacting to the AI video kind of in the corner. There was a reaction. We should bring that one back.

Host

我甚至生成过一只狗对我自己做出反应之类的。那很酷。但我们一直在尝试。我们试了很多不同的东西。甚至 UI 方面,早期乔伊有个想法,对于图像生成,就是一个简单的 UI,上面有你的联系人列表,然后你挑选哪些人出现在生成中。这有点道理,但当时并不明显。当然,事后看来很明显,这就是为什么它是“紫色代码”。但当你经历整个过程时,它并不明显。沿途有很多微小的决策,可能成就或毁掉你的产品。以 Instagram 为例,我记得和创始人之一米奇聊过,正方形照片是他们强加的一个限制。我想可能是因为他们不想处理宽高比的麻烦,而实际上这对它的成功至关重要。所以长话短说,我认为这个领域会变得非常、非常快,竞争非常、非常激烈。我相当有信心我们会保持领先。希望如此。而且我确实认为你需要在这个过程中拥抱人们,用创意工具赋予他们力量,而这正是我们在做的。

I even generated like a dog reacting to myself or something. It was cool. But we were trying things. We tried a lot of different stuff. Even the UI, it was early on Joey had this idea for image gen where it was just a simple UI where you had your contact list and it was like you just pick and choose people to be in the gen. That kind of makes sense, but it was not obvious. Of course, it's obvious in retrospect, that's why it's code purple. But it's not obvious when you go through this whole thing. There's a lot of micro decisions along the way that can make or break your product. With Instagram, I remember talking to Mikey, one of the founders there, and the square photos was a limitation that they just imposed. I think they probably because they didn't want to deal with aspect ratio nonsense and that was actually essential for its success. So long story short, I think this space is going to get very, very competitive very, very fast. I'm reasonably confident we're going to stay on top. Hope so. And I do think you need to embrace people in this whole process and empower them with creative tooling, and that's what we're doing.

产品方向与社区聚焦 Product direction and community focus

Host

看起来在产品方面你们一直在朝这个方向努力。我知道你们讨论过专注于社区以及产品中的其他方式。谈谈你们如何看待自己随着时间的推移进一步朝这个方向努力。

It seems like on the product side you're continually leaning into this. I know you guys have talked about focusing on communities and other ways in the product. Talk a little bit about how you see yourselves leaning into this even further over time.

Host

是的,这是一个冷启动的社交网络。不可能。有史以来最难的问题。

Yeah, it's a cold start social network. Impossible. Hardest problem ever.

Host

做起来很有趣。

Fun one to work on.

Host

是的,当然很有趣。尤其是面对一个全新的媒介。你会想,“好吧,让我想想这个。”所以我们并不确切知道事情会如何演变。我们给了它一个非常熟悉的形式。它看起来和其他全屏媒体播放器应用非常相似。它确实有不同的感觉,但我们有这个假设,而且我认为它正在被证明是正确的:和朋友一起玩更有趣。这个方面肯定体现在产品中。它在推荐系统中得到了强调,但我不知道它是否在所有可能的方式中都得到了充分强调。所以我看到我们随着时间的推移会更加朝这个方向努力。我认为公共信息流也非常重要,它甚至能给你灵感,比如“哦,实际上,沙克在这里。和沙克一起玩,或者对沙克关于玛丽莲·梦露的梗图进行即兴发挥,不管他今天在做什么,那都会很酷。”但我认为还有一些产品工作需要思考,而且我也很兴奋,这项技术能让你和朋友一起玩得更开心。可能还有一些我们没想到的东西,我们实际上可以朝那个方向努力。我不能确切地说它们是什么。

Yeah, it's fun for sure. Especially with a completely new medium. You're like, "Okay, let me think about this." So we didn't really know exactly how things would evolve. And we gave it a very familiar form factor. It looks very similar to other full screen media player apps. It definitely has a different feel, but we did have this hypothesis and I think it's kind of proving correct that this is much more fun with your friends. And that aspect definitely is in the product. It's emphasized in the recommender system, but I don't know if it's completely emphasized in all the ways that it could be. So I see us leaning into that a lot more over time. I think the public feed is also incredibly important and it's a way of even giving you inspiration to be like, "Oh, actually, well, Shaq's here. It'd be pretty cool to be with Shaq or have a riff on Shaq's Marilyn Monroe meme or whatever he's doing today." But I think there's some product work to be thought through and I'm also just excited about what this technology enables that is more fun with your friends. There's probably things that we haven't thought of yet that we can actually lean into. I can't say exactly what they are.

强调私信与群组动态 Emphasizing DMs and group dynamics

Bill Peebles

当然,我们越来越强调私信功能,因为我们认为那可以是一个非常神奇的瞬间。但你可以想象,小群体里发生的事情会非常令人兴奋。甚至大群体,我觉得也可能是一个很棒的切入点。当然,我们 OpenAI 在发布前的那个群组就非常酷。那是一个联系紧密、密度很高的大型网络,大家玩得很开心。我希望未来也能在产品中支持这种体验。

Certainly we're more emphasizing the DMs feature over time because we think that can be a very magical moment. But you can imagine things that happen in small groups being very exciting. Even large groups, I think, could be a very exciting angle. Certainly our OpenAI group before we launched was very cool. It's a very dense, connected, large network that people were just having a lot of fun with. I'd like to be able to support that in the product in the future as well.

Host

这很有意思。让我印象深刻的一点是,很明显,你设置信息流算法的方式基本上决定了生态系统中的很多行为,对吧?以及你在那里鼓励什么。

That's interesting. One thing I'm struck by is just clearly the way you put the feed algorithm basically determines a lot of this behavior in the ecosystem, right? And what you encourage there.

Bill Peebles

有句名言,我记不清是谁说的了,但大意是‘给我看激励,我就给你看结果’。这句话在大型推荐系统中再正确不过了。当然,我在 Instagram 工作时,有一个非常明确的决定,就是优先展示朋友的内容。我们选择了很多方式,不在主信息流中引入无关内容。我负责的是探索信息流,那是一个次要界面,但他们做出了那个非常明确的决定。随着时间的推移,这个原则有些偏离了,因为人们发帖变少了等等。但如果你看看现在的 Instagram 信息流,它有点奇怪。我觉得现在它相当令人不安。X 也是一样。我认为,因为我们大幅降低了创作门槛,我们很有机会重新强调这一点。很有趣的是,AI 视频居然能让你和朋友联系更紧密,这有点不寻常,但最热门的紧密联系网络是一个 AI 生成平台。但我完全支持这一点。

There's a famous quote. I'm trying to remember who did it, but it's like 'show me the incentives, I'll show you the outcome.' And this has never been more true than in these large recommender systems. Certainly when I worked on Instagram, there was a very explicit decision to prioritize your friends. We chose a lot of things to not introduce unconnected content in the main feed. I worked on the explore feed, which was a secondary surface, but they made that a very explicit decision. It's kind of drifted over time for the reason that people are posting less and all that sort of stuff. But if you look at your Instagram feed, it's kind of wacky. I'd say right now it's pretty disconcerting. Same with X. I think because we've brought this barrier to creation down so much, there's a good chance that we can actually emphasize that. It's very funny that AI videos are what bringing you more connected to your friends is a little bit unusual, but the hottest densely connected network is an AI gen. But I'm all here for it.

Host

是的。

Yeah.

Bill Peebles

我认为这在某种程度上是一个非常令人振奋的 AI 用例。让我印象深刻的一点是,它很有趣,而我觉得很多基于其他开放产品构建的东西都非常严肃。每个刚毕业的创始人想做的都是一个非常垂直的企业应用。真正被构建出来的消费产品其实不多。我觉得很多人都在尝试这个方向,比如如何让 AI 生成内容和人共处一室……我感觉 Character AI 早期也在即兴发挥一些类似的东西。我觉得你们基本上已经围绕这个方向做出了第一个成功的产品,这很酷。我的意思是,让世界看到这样的东西是件好事。

I think it's in some ways a very heartening use case for AI. One thing I was struck by is just it's fun in a way that I feel like so many of the things that have been built on top of other open products are just very serious. Every founder out of college wants to go build a very specific vertical enterprise app. There haven't been that many consumer products that have been built. I think a lot of people were playing around with this like how do you have AI generations and people in the same... I feel like Character AI was even riffing on some of this stuff in the early days. I feel like you guys have kind of nailed the first product around it and it's just cool. I mean, it's good cool for the world to see.

Bill Peebles

我的意思是,要获得你的 Cameo,并不是要经历一个疯狂的流程,比如音频挑战和摇头晃脑。有很多次,我和托马斯都觉得,我们完蛋了。人们绝对搞不定这个。

I mean, it wasn't like there's a whole crazy flow you have to go through to get your cameo, right, of like an audio challenge and moving your head around. There were many moments where Thomas and I were like, we're cooked. There's no way people are getting through this.

Host

我的意思是,你在说有趣的发布故事。我们当时在调试 Cameo,感觉这里面有点玄学,比如你要站在某个特定的柱子旁边。

I mean, you're talking about funny launch stories. We had because we were trying to dial in the cameos and we're like there, I don't know, there's a bit of voodoo involved in that whole thing where you stand in this particular poll.

Bill Peebles

哦,对。

Oh, yeah.

Host

时间必须是下午 4 点。

The time of the day is 4:00 p.m.

Bill Peebles

真的就是这样,我们都去了那个柱子那里。我们录了自己的 Cameo,结果不知怎么搞出了一个实际上是最优但根本不可能完成的流程——你得一边摇头一边说话。那其实恰恰相反。

Literally, it was this is that we all went to this poll. We had recorded cameos of ourselves and somehow we ended up at a flow which actually is the optimal flow but is impossible to do where you have to say the words while you move your head. That was actually the opposite.

Host

是啊。就像说一句话,同时这样摇头……太疯狂了。但我的意思是,它确实能生成稍微好一点的 Cameo。所以不管怎样,那是一个有趣的发布时刻。我们不得不把它调回来。就像,我觉得没人能搞定那个,因为我们自己都搞不定。

Yeah. It was like say a sentence move your head around like this like quick brown... It was insane. But I mean it would actually produce slightly better cameos. And so anyway that was a funny launch moment. We had to dial that back. Be like I don't think anyone's going to get through that one because we can't get through it.

Bill Peebles

是啊。这是不是个秘密专业技巧,如果我的背景看起来很有前景,我的 Cameo 就会更好?

Yeah. Is this a secret pro tip that my cameos would be better if my background looks very promising actually?

Host

在一个非常私密的地方打光,然后转动我的头。

Lighting a very private place and just roll my head in.

Bill Peebles

完全正确。

Exactly.

平衡消费者与专业创作者 Balancing consumer and prosumer creators

Host

我设想的一个矛盾是,因为你专注于创作者,所以你会立刻被卷入一个范围很广的复杂度焦点。我们之前聊过这个。有一种是纯消费型创作者,他们只想尽可能简单地混搭一些东西。然后还有这些专业消费者类型的用例,他们非常专业。你们显然已经引入了一些基本的编辑功能。你如何看待这个产品服务领域随时间的发展?

One tension I imagine you get dragged into immediately because you're focused on creators is there's just such a range of sophistication of focus. We were talking about this before. There's the kind of pure consumer type of creators that just want to be able to remix something as easily as possible. Then you have these prosumer type use cases where they're super sophisticated. You guys have obviously introduced some basic editing functionality. How do you think about this product service area over time?

Bill Peebles

我认为这个产品很酷的一点是,在最好的情况下,它非常民主化,让任何人都能创作,并逐步提升自己,成为传统意义上的专业创作者。如果你看到那些真正精通 Sora 的人生成的顶级作品,你可以直接混搭它,基本上可以直接访问其中的所有元素。你可以逐渐摸索如何掌握 Sora 的提示词技巧,如何设计你的 Cameo 等等。我认为其中一个重要部分是,继续让那些处于创意前沿的人,也就是这些专业消费者,能够用更好的工具获得赋能,这样他们就能不断突破。我们正在推出更多专门针对这个群体的功能。Storyboard 就是一个重要的功能。我们开始引入一些非常基础的编辑功能到应用中,比如本周早些时候上线的拼接功能。随着时间的推移,我希望每个人都能提升水平。我们会大力赋能顶级创作者去做他们的事情,但仅仅通过在信息流中看到这些内容,并拥有这些出色的混搭和编辑工具,每个人都能逐渐变成他们那样,这真的会让信息流成为一个令人惊叹的地方。一个最终状态是,更多的人能够发挥创意,而这也可以成为深入探索的入门途径,这其中有某种美好的东西。我想我小时候用 GarageBand 就有这种体验。这是一个很好的类比,它非常容易上手。最简单的操作就是拖拽循环片段。

I think one thing that's cool about this product is that at its best it's very democratizing in terms of who can create and who can kind of level up their game to become a traditionally pro creator. If you see any of these god-tier gens from people who have really mastered Sora, you can just remix it and have basically direct access to all the ingredients that went into that. You can sort of learn the ropes gradually about how you actually master prompting Sora, how you master designing your cameo, etc. I think an important part of this is continuing to let people who are at the frontier of creativity, these prosumers, continue to be empowered with even better tools, so they can keep pushing forward. We are launching more features that are really aimed towards that group specifically. Storyboard was a big one. We're starting to introduce very basic editing features into the app like stitching which went out earlier this week. Over time, my hope is everyone kind of levels up. We're going to really empower the top creators to do their thing, but just by virtue of seeing this stuff in the feed and having all these amazing remixing and editing tools, everyone can gradually turn into these people and it will really just make the feed an amazing place to be. There's something beautiful about an end state where more people can be creative and then also that can be a gateway drug for even going deeper. I think I had this experience growing up with GarageBand. It's a good analogy here where it was so accessible. At the simplest you could just drag loops.

降低创意门槛 Lowering the barrier to entry for creativity

Host

甚至不是真的坐下来演奏乐器,而是开始理解创作的要素是什么。你可以创作一些有趣的东西,但然后你可以深入,你会想,哦,现在我足够好奇了,想要买个 MIDI 键盘,学吉他,录音——这就像一扇入门药,但我们是通过降低准入门槛达到这一步的。我很好奇,我觉得现在每个应用开发者都在思考:我该围绕这些模型搭建什么样的脚手架,还是干脆去海滩度假两年,等模型自己变得更好?作为一个整体团队,你们如何看待围绕模型短板搭建的脚手架与模型自身改进之间的平衡?

Not even really sitting down and playing an instrument, but you're starting to get a sense of like, all right, what are the elements of creating? And you can actually create some interesting things, but then you can go deeper and you're like, oh, now now I'm actually curious enough to have the MIDI keyboard and like learn guitar and record and like it's a gateway drug, but you we got there by making the, you know, the barrier to entry much lower. And I'm curious, I feel like every app builder right now is is kind of thinking about like what do I build a scaffolding around these models versus like I just go to a beach for two years and like the models are going to get way better. How do you guys think about that as like a holistic team of like the the types of scaffolding you want to build around the shortcomings in the model versus ways it'll just get better and you know

Bill Peebles

是的。总的来说,我认为 OpenAI 的魔力很大程度上来自于拥有一个极其雄心勃勃的、与 AGI 对齐的研究路线图,并且无论发生什么——无论竞争对手发布什么,无论产品压力多大——都坚持执行。这当然也是我们对 Sora 的理念。我认为酷的是,随着这些模型越来越强大,我们发现了它们拥有的所有惊人能力,这让 Roan 和 Thomas 非常忙碌。所以像 Cameo、remixing 这样的功能,不仅仅是我们在研究上开创的特性,更是与产品团队那些出色的人合作的结果,他们共同探索这些模型能做的所有疯狂事情。

Yeah. In general, I think a lot of OpenAI's magic comes from just having this incredibly ambitious AGI aligned research roadmap and sticking to it no matter what happens, no matter what the competitors release, no matter what the product pressure is. Um, and that's certainly our philosophy on Sora. You know, I think what's cool is as these models get more and more powerful, we kind of discover all these amazing capabilities that they have and that, you know, keeps like Roan and Thomas very busy. Um, so you know something like Cameo, something like remixing, right, is really uh not just uh a feature that we pioneer on research, but it's really a joint collaboration with the work the amazing guys on product are doing to figure out all the crazy things you can do with these models.

Host

我确实认为这需要你用非常创造性的视角来看待它,或者愿意失败。比如,试试完全由 AI 生成、没有 Cameo 的内容——其实没那么酷。但我可以想象,在游戏和其他领域,即使是今天的语言模型和视频模型也能以非常有趣的方式支持很多应用。所以我认为只需要一点跳出框框的思考,而不是试图以任何方式复制 OpenAI 的路线图。而是思考:什么样的变体才真正有意义?我一直在说消费端,不是企业端——企业端当然有成千上万的应用——但消费端的东西就是:什么有创意?什么新鲜?拥抱它。你能做的远不止这些。有控制力,有各种令人兴奋的东西。

I do think it requires a very creative lens on what you do with it. Um, or a willingness to fail, you know, like let's try the AI, you know, completely AI feed with no cameos. It's actually not that cool. But I I can imagine a lot of things in gaming and other other surfaces that even today's LMS and video models could support in a very interesting way. And so I think it's just a little bit of thinking outside the box and not trying to mirror the exact road map of Open AI in any way. Uh but being like okay like what's a spin on this that actually makes sense. I'm always consumer build not talking about enterprise of which there are a thousand for these models of course but like the consumer things is just like what's creative? What's new? Let's embrace that. There's way more you can do with this. There's control. There's all kinds of exciting things. Yeah.

Bill Peebles

是的。这也是我们随着这一代模型推出 Sora API 的重要原因。正如 Thomas 所暗示的,你可以用这些东西构建太多东西了。而且 Sora 团队非常小。

Yeah. This is a big part of the reason we started launching a Sora API as well with this generation of models. There is so much stuff as Thomas is alluding to that you can build with these things. And we're a very small team on Sora.

Host

小得惊人。我是说,

Shockingly small. I mean,

Bill Peebles

发布时不到 20 人,现在大概 50 人左右。

shockingly um you were under 20 when you released it and now we're like 50 or so.

Host

我们大概有 9 到 10 个研究人员在 Sora 上。产品团队大概有多少?

We're at like So there's like roughly nine or 10 researchers on Sora. I think product is at what like

Bill Peebles

不到 20 人。然后我们还有一个系统团队,大约 13 人左右。所以总共大约 40 人。相当小。所以那些联系我们创作者,或者那些想构建新应用的创业者,现在都可以通过 Sora API 来实现。

under 20ish. Yeah. Um and then we have a systems team which is about like 13 or so folks. So it's it's like 40ish in total. It's pretty small. And so you know all of uh you know these creators who reach out or uh you know all these uh very entrepreneurial folks who want to build these new applications they now have the ability to do that with the Sora API.

Host

你们在 Sora API 上已经看到什么酷的东西了吗?或者人们对它能做什么让你们感到兴奋?

Have you seen anything cool already on the Sora API or like what gets you guys excited about uh what people can do?

Bill Peebles

Mattel 做了一些很酷的东西。他们一直在原型设计

Mattel has done some cool stuff. So they've been prototyping

Host

比如芭比 Mattel?

like Barbie Mattel.

Bill Peebles

是的,他们一直在用 Sora 原型设计新玩具,这看起来非常酷。我认为随着时间的推移,应用会变得越来越复杂。API 发布才三周,但已经有人在做一些相当疯狂的事情了。我看到一些 CAD 到可视化的流程,人们把 CAD 文件转换成 Sora 能理解的描述,以便可视化零件。我一开始觉得这似乎不太精确,但他们解释了为什么这实际上很关键,而且是缺失的东西。我当时想,哇,这太棒了。

Yeah they've been prototyping new toys with Sora which is just super cool to see. Um, you know, this is I think you know over time the applications will just get increasingly sophisticated. It's been I think 3 weeks since we launched the API but yeah already people are doing some pretty wacky stuff. I've seen some like CAD CAD to visualization pipelines where people are taking like a CAD file trans translating that to some caption that Sora could understand so they could visualize their parts which I'm like that doesn't seem very precise and they're like explained why this is actually critical and something missing and they're I'm like whoa that's yeah amazing

AI视频生成进展 Progress in AI video generation

Host

也许退一步,谈谈视频模型这边。

you know maybe take a step back and talking about the video model side like

Bill Peebles

让我为听众们梳理一下背景:过去几年 AI 视频领域取得了哪些进展,有哪些重要的里程碑。我想说,有一段时间视频领域基本没有进展,大部分进展都在图像生成方面。图像生成领域最早的重要论文之一是 OpenAI 几年前发布的 DALL-E 1。那是 Aditya 的早期工作,也是我们第一次看到语言模型中那种突破性的通用能力迁移到视觉生成领域。在那之前,模型非常小众,只能建模非常狭窄的分布,比如人脸。但那时还没有一个时刻能清楚地表明这些模型可以建模互联网上的所有像素。我认为从那时起,方向就明确了。又花了几年时间才真正完全控制图像生成。所以有了 DALL-E 2、DALL-E 3。我们在 2023 年初开始研究 Sora,与 DALL-E 3 的开发同期进行。当时我们拥有所有必要的要素,取得了重大突破。我们开始理解 Scaling 的工作原理。扩散模型在公式和架构上都变得更加原理化。Sora 1 就像是视频领域的 GPT-1 时刻。那是第一次你可以考虑做超过 1 秒的高分辨率一致生成。用那个模型我们勉强能做到 60 秒。从那以后,我们一直在努力推动智能和可用性的前沿。很长一段时间我们都在想 Sora 2 相当于 GPT 的哪个版本。现在很清楚,在很多方面它相当于 GPT 3.5,无论是能力上的突破。这是唯一一个能完成奥运体操动作而不会让肢体乱飞的模型。

but let me just kind of contextualize for our listeners like how has the AI video space progressed over the past few years maybe some like the different milestones that matter to So I'd say for a while uh basically there was no progress in video and a lot of it was really on the image generation side. So one of the most important early papers in image generation um was DALL-E 1 which came out of uh OpenAI several years back. Um that was like early work that Aditya did and that was really the first time we kind of saw this you know breakthrough general purpose capability that we were starting to see in the LMs come to the visual generation side. Before that point, you know, there were models that were very niche in terms of modeling very specific narrow distributions like human faces. Um, but there was really not a moment where it was clear these things could model kind of all the pixels that exist on the internet. Uh, I think from that point on it became clear where things were headed. It took a few years to really get image generation under full control. So, you know, DALL-E 2, DALL-E 3. Um, we started working on Sora in early 2023. It was co-developed at the time that DALL-E 3 was being worked on. and we kind of had all of the necessary ingredients like make a big breakthrough at that point in time. We started to started to understand how scaling works. Uh diffusion models were starting to become much more principled uh both in terms of the actual formulation of diffusion as well as the architecture of them. And Sora 1 was really like the GPT-1 moment for video. Um that was the first time you could even think about doing kind of high-res consistent generation uh for anything more than 1 second. Um we could do 60 seconds barely uh with that model. And since then, you know, we've really just been trying to push the frontier of intelligence and usability. So we for a long time we were actually trying to figure out what number of GPT is Sora 2. Yeah, I think it's we're very clear now it's GPT 3.5 in many ways both in terms of the breakthrough in capabilities. This is the only model that can you know do an Olympics gymnastic routine and like not have things just like go crazy,

Host

对吧?我喜欢这个评估——大家都在 Twitter 上发奥运提示词,结果看到肢体到处飞。我们被狠狠吐槽了。所以很明显,这既体现在模型智能的提升上,也体现在可用性方面。GPT 3.5 带来了 ChatGPT。

right? I love that that's one of the eval uh everyone just like went off on Twitter with like all these Olympics proms. He saw limbs flying everywhere. Yeah, we got we got roasted. Um, so clearly, you know, it it's justified by the model intelligence boost, but also in terms of the usability side, right? GPT 3.5 ushered in chat GPT.

从GPT-1到3.5的惊喜 Surprise at Progress from GPT-1 to 3.5

Host

从 GPT-1 到 3.5 的飞跃让你感到惊讶吗?还是说圈内人早就预料到,只要继续 Scaling 就会发生?

Did the jump from GPT-1 to 3.5 surprise you? Or did people in the know expect it if we kept scaling?

Bill Peebles

我们知道我们发现了有价值的东西。看到模型理解物理的能力快速提升——比如体操动作能成功,或者杯子掉下来真的摔碎——非常直观。我们知道进展不会像语言模型那么慢;我们正在加速时间线。我们不确定能加速多少,但 3.5 最终达到了预期的高端。

We knew we were onto something. Seeing rapid improvement in how the model understands physics—like gymnastics routines working or a glass dropping and shattering—was very visceral. We knew the ramp wouldn't be as slow as for LMs; we were accelerating timelines. We weren't sure by how much, but 3.5 ended up on the upper end of expectations.

下一突破:长时模拟 Next Breakthrough: Long-Duration Simulation

Host

从 3.5 到 4 的飞跃是什么?有哪些未解决的问题需要改进?

What is the jump from 3.5 to 4? What unsolved problems need improvement?

Bill Peebles

视频的下一个突破性能力是模拟持续数小时甚至更长的过程。我们很兴奋 Sora 能用于知识工作、生物学、物理学研究。模拟湿实验室需要让模型运行数天、数周、数年,解决生成建模中的许多基本问题。

The next breakthrough capability in video is simulation of processes lasting hours or longer. We're excited about Sora being used for knowledge work, biology, physics research. Simulating a wet lab requires running models for days, weeks, years, solving fundamental problems in generative modeling.

视频模型在机器人学与科学中的角色 Role of Video Models in Robotics and Science

Host

仅靠模拟数据在机器人领域能走多远,尤其是操作任务?我们是否需要大规模的真实世界数据收集?

How far can simulation data alone get us in robotics, especially for manipulation? Will we need massive real-world data collection?

Bill Peebles

视频模型将发挥关键作用。机器人领域缺乏用于预训练的大规模轨迹数据集,但视频模型深刻理解局部运动和灵巧操作。我们对复用像 Sora 这样的模型非常乐观。我相信基于早期视频模型的旧研究方向会随着基础模型智能的提升而取得成功。

Video models will be instrumental. Robotics lacks large trajectory datasets for pre-training, but video models deeply understand local motion and dexterity. We're bullish on repurposing models like Sora. I believe older research lines based on early video models will succeed with increases in base model intelligence.

Host

在生物、机器人、材料科学领域工作,是否都需要最先进的视频模型?

Will you need a state-of-the-art video model to work in bio, robotics, material sciences?

Bill Peebles

越来越是这样。

Increasingly, yes.

世界模拟目标与里程碑 World Simulation Goal and Milestones

Host

所以下一个前沿是更长的视频和物理保真度?

So the next frontiers are longer videos and physical fidelity?

Bill Peebles

是的,这是为了追求世界模拟的目标。我们希望 Sora 深刻理解现实的每一个细节,并超越娱乐用途。娱乐很棒,但这只是第一阶段;实用性将飞速增长。

Yes, it's in pursuit of the world simulation goal. We want Sora to deeply understand every bit of reality and be useful beyond entertainment. Entertainment is great, but this is phase one; usefulness will skyrocket.

Host

你心中有没有一个评估指标能标志着一个里程碑?

Is there an eval you have in mind that would mark a milestone?

Bill Peebles

通过视频模型模拟某个现象带来的第一个科学突破将是一个疯狂的里程碑。可能与经典物理有关,比如湍流。那将标志着一个新时代的开始。

The first scientific breakthrough from simulating a phenomenon with video models will be an insane milestone. It could be related to classical physics, like turbulence. That will mark the beginning of a new era.

Host

对时间有什么预测吗?

Any prediction on when that might be?

Bill Peebles

如果到 2028 年初我们还没有取得这样的突破,我会感到震惊。

I'd be shocked if we aren't making breakthroughs like this by early 2028.

视频缩放定律与新突破 Scaling Laws and New Breakthroughs in Video

Host

在 LLM 领域,进步来自缩放定律、数据、算法突破和更多 GPU。视频领域也类似吗?

In the LLM world, progress comes from scaling laws, data, algorithmic breakthroughs, and more GPUs. Is it similar for video?

Bill Peebles

进步有很多维度;Scaling 是其中之一。但对于长达数年的模拟运行,我们可能需要新的突破——不能直接移植现有技术。构建一个能记住每个细节的替代现实可能需要新的东西。有很多未开发的领域值得探索。

There are many axes for progress; scaling is a big one. But for years-long simulation rollouts, we may need new breakthroughs—you can't just port existing techniques. Building an alternate reality that remembers every detail might require something new. Lots of green pastures to explore.

评估视频模型 Evaluating Video Models

Host

你们现在如何对视频模型进行评估?比如,怎么知道它们在变好?在产品端你们又是怎么做的?有没有一些惯用的方法?

How do you do evals on video models now? Like what, how do you know when they're getting better? And how do you guys do on the product side? Do you have any go-to things?

Bill Peebles

我认为这是我们已经成熟很多的领域,尤其是从 Sora 1 到 Sora 2。我们学会了如何利用真实用例进行高杠杆的有效评估。在发布之前,对用例的信心让我们能够构建基础的产品评估。比如,“嘿,我们正在改变模型或 XYZ 中的某些部分,让我们用 Sora 2 运行 Sora 1 的顶级提示,看看差异。”现在我们有实际的生产用例,比如 Cameo 就是一个很好的例子,每次我们做出改变,都想了解它如何影响这些核心用例之一。

This is an area where we've matured a lot, I think, especially from Sora 1 to Sora 2. Just learning how high-leverage good evals with real use cases are. Before we launched, having conviction in the use cases lets us build basic product evals. For example, 'Hey, we're changing the model or some part of the stack in XYZ, let's run the top Sora 1 prompt through Sora 2 and get a sense of the difference.' Now that we actually have production use cases, like Cameo is a great one, anytime we make a change we want to understand how it impacts one of these core use cases.

Host

是的。

Yeah.

个性化vs单一模型 Personalization vs. Single Model

Host

我很好奇,对于视频模型,显然你可以想象,如果平均所有人的偏好,你可能会得到一个没有个性的模型,对谁都不完美。我想知道,随着时间的推移,你认为最终会出现一堆具有不同美学和不同用例的模型,还是它会收敛到一个你可以根据自己的喜好来引导的单一模型?

I'm curious with video models, obviously you could imagine if you averaged out the preferences of everyone, you may end up with a personality-less model that is perfect to no one. I wonder over time, do you think there end up being a bunch of different models with different aesthetics and different use cases, or does it kind of converge on a single model that you can steer to your liking?

Bill Peebles

我认为这里实际上有一个与推荐系统的类比,非常贴切。当人们想到推荐系统时,不一定每个人都经常思考他们的信息流,但假设让你设计一个信息流。你首先想到的可能是,“哦,让我按流行度排序或按类似的东西排序。”然后你会发现这正是问题所在:这种回归到平均的情况,实际上对谁都不有趣,因为它只是一个全球流行的趋势,并不符合你个人的偏好。所以神奇之处在于,你以不同形式引入个性化。有很多不同的方法可以做到这一点,但最明显的一个是查看你的历史记录,看看你接下来要做什么。这确实改变了信息流中的感觉和相关性结果。我无法想象我们在几乎所有地方不会出现类似的现象。我的意思是,我们在 ChatGPT 中看到很多关于个性化非常重要的事情,比如模型与你交谈的方式。我不一定直接谈论视频模型,但在推荐系统中,你很快就会发现你不希望为每个人定制模型,对吧?那是一个基础设施的噩梦。而且从很多方面来说,它也不符合追求 AGI 的目标。你可以做到,但它并不真正可扩展。而且也很难推理。更好的方法是利用群体的智慧,找到人与人之间的类比。你进行协同过滤,规模以一种美妙的方式让你受益:就像我知道这个人过去喜欢过类似的视频,而我和那个人相似,所以因此我也很可能会喜欢这个视频,或者从这个视频中获得创作灵感。

I think there's actually an analogy here to recommender systems. It fits pretty nicely. When people think of recommender systems, not that everyone necessarily thinks through their feed very often, but let's just say it tasks you with designing a feed. Your first thing you might think of is just, 'Oh, let me sort by popularity or sort by something like that.' And you find that exactly is the problem: this regression to the mean kind of thing where it's actually interesting to nobody because it's just a trend that is only globally popular, which is not suiting your own personal preference. So what the magic is, you introduce personalization in different forms. There are lots of different ways of doing this, but the obvious one is look at your history and see what you're going to do next. That really does change the feeling and the results of relevance in your feed. I can't imagine that we wouldn't have a similar phenomenon basically everywhere. I mean, we're seeing a lot in ChatGPT around personalization being a very important thing, the way the model talks to you. I wouldn't necessarily speak to video models directly, but in recommender systems you learn very quickly that you don't want bespoke models for everybody, right? It's an infrastructural nightmare. It's also not really in the pursuit of AGI in many ways. You can kind of do it, but it doesn't really scale. It also gets very hard to reason about. Much better is to leverage the wisdom of the crowds where you find analogies between people. You do collaborative filtering, and it helps actually the scale benefits you in this beautiful way: it's like I know that this person has liked videos that are similar in the past, and I'm similar to that person, so now therefore I'm going to very likely enjoy this video as well or be inspired to be creative from this video.

Host

是的。我的意思是,我认为以人类的方式感知世界的模型,如果你以那种方式衡量智能,那么基础模型中内置的多样性和对不同风格行为的理解是这种智能的重要组成部分。对我来说,这是 Sora 2 最疯狂的事情,不是最疯狂的,但最令人惊叹的事情之一。我认为今天几乎所有其他视频模型都存在模式崩溃。即使是那些擅长电影内容和物理的模型。Sora 感觉它既擅长那个,也擅长像门铃摄像头那样的镜头。它擅长像采访、播客这样的内容。它擅长所有这些不同种类的东西,感觉非常动漫化,风格迥异。模型的范围令人惊叹,我只希望我们继续朝这个方向努力。

Yeah. I mean, I think models that perceive the world the way humans do, like if you measure intelligence in that way, diversity and understanding different stylistic behavior baked into the base model is an important part of that intelligence. And that was the craziest thing about Sora 2 to me, not the craziest thing, but one of the most amazing things. There's definitely mode collapse in I think almost all the other video models out there today. Even the ones that are great at cinematic content and physics. Sora feels like it's great at that and great at like a ring doorbell footage. It's great at like an interview, like a podcast. It's great at all these different kinds of things that feel very anime, very stylistically different. The range of the model is amazing, and I only anticipate that we lean into that.

Bill Peebles

是的。

Yeah.

成本降低轨迹 Cost Reduction Trajectory

Host

我的意思是,你们显然在语言模型方面宠坏了人们,默认的期望是你们发布的任何东西在 6 到 12 个月内都会便宜 100 倍。我们应该在视频方面期待类似的事情吗?

I mean you guys have obviously spoiled people on the LM side where it's just like the default expectation is whatever you release is like 100 times cheaper in 6 to 12 months. Should we expect a similar thing on the video side?

Bill Peebles

100%。即使你回到 2024 年 2 月,当我们第一次向世界展示 Sora 1 时,那个模型采样一个 720p 的短视频需要大约 50 美元的算力。如果你看看我们 Sora 2 的 API 定价,与之相比只是几分钱。所以我们已经看到成本降低了几个数量级,同时模型智能大幅提升,这一趋势肯定会继续。

100%. And even if you rewind the clock to February 2024 when we showed off Sora 1 to the world for the first time, that model cost like $50 in compute to sample a 720p short video. And if you look at our API pricing for Sora 2, it's like cents on the dollar compared to that. So already we're seeing orders of magnitude decreases in cost with vastly higher model intelligence with this release, and that trend will definitely continue.

变现与生态系统 Monetization and Ecosystem

Host

你想在这个播客上发布重大新闻。所以我们开始前 30 分钟,你发推说你要推出一些定价,我认为这完全合理。我觉得当你每天有 30 次免费使用,然后开始为它们支付一点费用时,这很自然。而且从我所看到的,互联网并没有暴动。这似乎是货币化的自然第一步。我知道你们之前谈过,既要为运行这些模型的推理成本进行货币化,也要通过货币化来想办法激励版权持有者和各种人参与进来。

You wanted to break big news on this podcast. So 30 minutes before we started, you tweeted that you were going to introduce some pricing, which I guess is entirely reasonable. I think when you get 30 a day and then you have to start paying a little bit for them. And I don't think the internet rioted from what I saw. That seems like a natural first step in monetizing. I know you guys have talked before about both monetization for the inference cost of running these models, but also monetization to figure out how to incentivize rights holders and all sorts of folks to get involved.

Bill Peebles

是的,我的意思是,我们真的想创建一个让每个人都受益的生态系统。比如我们需要支付 Sora 的 GPU 账单。我们希望即使是这个平台上新兴的创作者,他们不一定在 TikTok 或 Instagram 上有粉丝,我们也希望他们有一条成功和赚钱的道路。然后我们希望拥有令人难以置信的角色库的版权持有者也能受益,很多人会喜欢使用这些角色。所以当我们考虑最初的货币化方式时,我们每天都在学习很多。我们想慢慢来,确保我们在最重要的类别上打勾,即为我们生态系统中的所有这些不同的人提供货币化的途径。我们确实将积分视为实现这一目标的主要方式,而不必过度承诺一个我们不知道能否长期支持或最终需要关闭的模型。所以我们不认为这一定是最终货币化 Sora 的方式。我们对此非常开放。

Yeah, I mean we really want to create an ecosystem where everyone is benefiting. Like we need to pay the GPU bills for Sora. We want even new creators who are coming up on this platform, they didn't necessarily have a following on TikTok or IG, we want them to have a path to success and making money. And then we want the rights holders who have this incredible library of characters that so many people will just love using to also benefit. So when we were thinking about the initial ways to monetize, we're going into this learning a lot every day. We want to kind of take it slow, and we want to make sure we're checking the boxes on the most important categories here, which is giving a path to monetization to all of these different people within our ecosystem. And we really viewed credits as being kind of the primary way to do that, without necessarily overcommitting to a model that we don't know we can support long term or just need to pull the plug on eventually. So we don't think this will necessarily be the final way that we ultimately monetize Sora. We're really open-minded in this.

透明度与创作者变现 Transparency and Creator Monetization

Host

我们尽量保持透明,做出这些决定并听取各方的意见,因为我们希望找到一个平衡点,让 OpenAI 不是唯一赚钱的一方,而是所有人都能受益。我认为这对平台的长期成功非常重要。这应该能成为某人的全职工作。如果创作者想在 Sora 上走红,我们应该有办法让他们实现这一点。你们还考虑过其他定价模式吗?

We're trying to be as transparent as possible. We make these decisions and hear what everybody has to say because we want to land this in a spot where not only OpenAI is making money, but everybody benefits. I think that's really important to the long-term success of the platform. This should be somebody's full-time job. If they're a creator and want to go viral on Sora, we should have a way to make that possible. Any other pricing models you've thought about?

Bill Peebles

短期内没什么特别具体的。但回到如何利用生成式视频重新思考品牌和赞助的问题上。现在,如果你刷 Instagram,会看到视频广告。但如果创作者允许视频中的所有无生命物体都变成特定品牌,并把它们拍卖给品牌呢?有一些很疯狂的想法。

Nothing super tangible in the short term. But going back to how you can rethink branding and sponsorship with generative video. Right now, if you're scrolling Instagram, you get a video ad. But what if a creator is okay with all inanimate objects in their video being certain brands, and they auction that off to brands? There are some wacky ideas out there.

客串与触达个人体验 Personal Experience with Cameo and Reach

Host

这个平台的另一个有趣之处在于触达范围。我作为最早的用户之一公开了我的 cameo。现在我有大约 17,000 次 cameo 出场。如果算总观看次数,我这辈子可能从未有过这么大的传播量。而且我是个无名小卒。我的粉丝数快追上 Twitter 了。这太疯狂了。再乘以 cameo 次数,这种触达在其他平台上几乎不可能,因为你必须自己创作。我喜欢每天看到自己做了什么——我更新 cameo 指令。我也在看别人在做什么。这非常有趣。我觉得我们还没有一个很好的类比来描述这种格式。有很多新媒介的东西,但有趣的是,触达不仅仅取决于你发布的内容。

Another interesting thing about this platform is the reach. I made my cameo public as one of the first early adopters. I now have about 17,000 cameo appearances. If I sum the view count, I've probably never had more distribution in my entire life. And I'm a nobody. My follower count is creeping up to my Twitter one. It's crazy. Multiply that by the cameo count, and there's an impossible level of reach on other platforms because you have to create it yourself. I like seeing what I do every day—I change my cameo instructions. I'm also seeing what other people are doing. It's a very interesting thing. I don't think we have a great analogy for this format. There's a lot of new medium things, but it's interesting to think that reach is not just what you post.

全球使用与文化差异 Global Usage and Cultural Variations

Bill Peebles

这个产品的一个酷点是它是全球性的。我想知道你是否看到不同地区的不同用法。

One cool thing about the product is that it's global. I wonder if you've seen different variations across geographies.

Host

我们昨晚刚在几个东南亚国家上线。新鲜出炉。最初我们在美国、加拿大上线,然后扩展到韩国和日本。创作风格确实有很大不同。对 cameo 和角色 cameo 的接受度不同。这很鼓舞人心。有时人们抱怨跨文化内容,但我喜欢。一切都各有风味。我们都喜欢和关注的日本创作者——非常唯美,与美国的一切感觉都不同。我相信还会有更多。我喜欢从生成内容中学习。我看到有人谈论多伦多口音。我来自多伦多北部,之前都不知道有这种口音。我立刻打开维基百科,发现了一篇长文。我在 TikTok 上也有类似体验。我学到了很多东西——关于自己的心理现象、依恋理论。我认为我们还会看到更多。跨文化的东西就在那里。当你看到混剪链时,看到不同国家参与其中,各有风味,很有趣。说唱也略有不同。

We just launched to some Southeast Asian countries last night. Hot off the press. We originally launched to the US, Canada, then Korea and Japan. There's definitely a huge different flavor of creation. Comfortableness with cameos, character cameos. It's inspiring. Sometimes people complain about cross-cultural content, but I love it. Everything has a different flavor. The Japanese creators we all love and follow—it's very aesthetic, a different feel from anything in the US. I'm sure there's more to come. I like learning from the gens. I saw someone talking about the Toronto accent. I'm from north of Toronto and didn't know there was one. I immediately opened Wikipedia and found a huge article on it. That was my experience with TikTok too. I learn so many things—psychological things about myself, attachment theory. I think we'll see a lot more of that. Cross-cultural stuff is there. When you see a remix chain, it's fun to see all the different countries participating, each with a different flavor. The rap is slightly different.

Bill Peebles

我需要一个翻译按钮,但已经很接近了。

I need a translate button, but it's so close.

Host

我喜欢最新动态。那是我最喜欢的窗口,能看到人们在应用里做什么。就是实时发生的事情。接着汤姆说的学习,就是了解人们想看到自己和朋友出现在什么场景里。人们拿来玩梗的东西太有趣了。对一个人来说很平常,但如果你进入那种关注人们关心和笑点的模式,就会发现很多好笑的东西让人着迷。我今天和一个用户聊天,她痴迷于起重机。她说:‘我就是喜欢想象自己站在起重机顶上。’

The latest feed I love. It's my favorite spotlight into what people are doing on the app. It's just the stream of what's happening. Riffing off what Tom said about learning things, just learning what scenes people want to see themselves and their friends in. It's so interesting the kind of things people are memeing about. It's mundane to one person, but if you're in that mode of what people care and laugh about, there are so many hilarious things people are obsessed with. I was talking to a user today who's obsessed with cranes. She said, 'I just love visualizing myself on top of cranes.'

审核与安全 Moderation and Safety

Bill Peebles

我觉得你们在审核方面考虑得非常周到。既鼓励有趣的内容,又处理得很好。是迭代改进还是一开始就做对了?说说看。

I think you guys have been super thoughtful on the moderation side. Encouraging fun content but also navigating those waters well. Was it iterative or did you nail it from step one? Talk about that.

Host

天哪,我们熬了无数个不眠之夜才走到今天,而且还有很多工作要做。感谢 Sora 团队和 OpenAI 负责安全、审核系统和模型的优秀同事们。我们的安全栈里有推理模型,非常棒。这是用技术改进产品的好例子。现在我们在审核和安全方面能用更小的团队做更多事。但谢谢你认为我们做得不错。Twitter 上的人可不是这么说的。

Oh my god, we had long sleepless nights getting to where we are now, and still a lot of work. Shout out to the amazing people working on safety on the Sora team and at OpenAI on moderation systems and models. We have reasoning models in our safety stacks that are amazing. It's a great example of using our technology to make products better. We're able to do a lot more with a smaller team on the moderation and safety front. But thank you for saying it's in a good place. Not what people on Twitter tell us.

内容审核挑战 Content moderation challenges

Host

但 Twitter 上有人对某些事感到不满。我无法……内容审核的失败是可以理解的,令人沮丧,因为我们试图在用户自由和屏蔽所有不良内容之间走钢丝。我们每天都在进步。我们正处于新领域,尤其是客串功能。我们希望非常敏感地让你感到自己掌控一切,并在这个网络上感到舒适。因此,我们在护栏上投入了很多思考。

But Twitter angry about something. I can't... there's the content moderation failures are understandably frustrating because we're trying to tow this line between user freedom but blocking all the bad stuff. And we're getting better every day. We're in new territory especially with cameos. We want to be really sensitive to you feeling like you're in control and comfortable on this network. So with that comes a lot of thought in the guardrails.

Host

是的。你曾提到这是一个新的社交平台。我想 Sam 非常公开地谈到过,这对 OpenAI 来说是一个有趣的探索领域。你提到过这可能随着时间的推移与 ChatGPT 整合。你如何看待 OpenAI 正在研究的其他方面,比如有一天与我们融合?

Yeah. You've kind of talked about this being a new social platform. I think Sam's spoken very publicly about that being interesting surface area for OpenAI to explore. You talked about maybe this integrating with ChatGPT over time. How do you see other angles of things being worked on in OpenAI like one day coalescing with us?

Bill Peebles

我是说 Bill 提到过超越娱乐。ChatGPT 就像你的超级助手。早期我们就是这么叫它的。为什么它不能用真正有用的视频来回应呢?我认为一旦我们加入这类功能,那将非常棒。但没错,我们的产品可以以各种方式互动,世界尽在掌握。我是说,我们刚刚发布了一个浏览器。我不知道,你可以想象使用浏览器时,旁边有一个小视频助手之类的东西在跟你说话,你的智能体帮你订这张机票。我听说过一些疯狂古怪的想法。

I mean Bill mentioned going beyond entertainment. ChatGPT is like your super assistant. That's literally what we called it early days. Why shouldn't it be able to respond with a really helpful video? I think that'll be really killer once we get that kind of stuff in. But yeah, it's like the world's our oyster with ways that all of our products could interact. I mean, we just released a browser. I don't know, you can imagine using a browser and having a little video assistant or something talking to you on the side, your agent help me book this flight. There's crazy wacky ideas I've heard out there.

Host

是的。而且我认为很多东西是相互促进的。比如推理模型,在追求 AGI 的过程中非常合理。我的第一个用例不会是审核堆栈。等等,那实际上是一个完美的用例。所以很多东西是相互促进的。生态系统本身,我确实觉得 ChatGPT 有点不同,有点神圣。这并不意味着它不能随时间改变。我们肯定可以在那里做一些事情,但它往往是一个更注重实用性的用例。将娱乐融入实用性用例并不总是有效。所以我不认为只是把它塞进去就能神奇地工作。我认为我们需要非常深思熟虑地去做。但当然,作为一个实用性用例,这些模型将发现湍流建模的新模式。所以显然有实用性。

Yeah. And I think a lot of this stuff builds off each other. Say the reasoning models which in the pursuit of AGI is very justified. My first use case wouldn't be a moderation stack. Wait a second, that actually is a perfect use case. So a lot of this stuff builds on each other. The ecosystem itself, I do think ChatGPT feels a little different, a little bit sacred. It doesn't mean it can't change over time. There's definitely things that we can do there, but it tends to be a more utility driven use case. Mixing entertainment in a utility driven use case doesn't always work. So I don't think it's just like jam it in there and it will magically work. I think we need to do that very thoughtfully. But certainly as a utility use case, these models are going to discover new modes of turbulence modeling. So there's clearly utility.

Host

我听说 2028 年,甚至不是第二季度。

I heard 2028 the early not even by the second.

Bill Peebles

你应该可以问 ChatGPT 这个问题。而视频作为一种模态,你在 YouTube 上看到,一半的视频都是教程视频。这是我们希望建模的东西。但就我而言,这是我的马桶,修好它,弄清楚如何调整并解决这个问题。所以我认为这可以非常自然地连接起来。

You should be able to ask ChatGPT about that. And video as a modality, you see it on YouTube like half the videos on YouTube are how-to videos. That's something that we would want to see modeled. But in my case, here's my toilet, fix it, figure out how I adjust it and fix this problem. So I think that can be very naturally connected.

模型学习物理的方式 How models learn physics

Host

是的。在我们进入快速问答环节之前,我想问你一个厚脸皮的问题:解释这些模型如何学习物理学的最简单方法是什么?

Yeah. Before we move to our quick fire round, I did want to ask you a shameless question: what's the simplest way to explain how these models learn physics?

Bill Peebles

这是个好问题。在高层次上,这些模型总是在做预测任务。以扩散模型为例,你拿到一个视频,我们人为地添加了大量噪声,神经网络的目标是预测被噪声掩盖的底层信号。语言模型也做预测任务:你根据所有先前的词元预测下一个词元。我们实际上认为语言模型和 Sora 都学习世界模型。它们学习不同风格的世界模型。但根本上,如果你在做这个预测任务,无论是预测信号还是预测下一个词元,你都需要对世界的动态有所理解。所以对于语言模型,如果我生成一首押韵诗或一首诗,拥有关于诗歌结构、语言元素的内在知识非常有用,这些知识是从数据中习得的。因为如果我没有理解这些东西,就很难预测下一个词元。视频也是如此。如果我是视频模型,提示是一个人在打篮球,我最好已经学会了篮球如何运球、光线如何折射穿过场景,以及所有这些物理的小组成部分,才能描绘出整个场景的更大画面。所以这基本上是从大量算力和大量数据中涌现出来的属性。这在今天是很老生常谈的话了。大家都懂。更好的教训有效。但根本上,如果你没有这些关于世界如何运作的内在模型,你永远会比拥有它们的模型更差,从而有更高的损失。这就是我们如何看待优化压力,它使物理从大规模视频预训练中涌现出来。

It's a great question. At a high level, these models are always doing prediction tasks. In the case of diffusion models, you're getting a video which we've artificially added a lot of noise to, and the goal of the neural network is to predict the underlying signal that we obscured with all that noise. LMs also do a prediction task: you're predicting the next token conditioned on all the prior tokens. We think actually both LMs and Sora learn world models. They learn different flavors of world models. But fundamentally, if you're doing this prediction task, whether it's predicting signal or predicting the next token, you need to develop some understanding of the dynamics of the world. So in the case of LMs, if I'm generating a rhyme or a poem, it's very useful for me to have some internal knowledge about the structure of poems, about the linguistic elements that I pick up from data. Because if I haven't grokked these things, it's very hard for me to predict the next token. The analogy holds for video as well. If I'm a video model and the prompt is a guy playing basketball, I sure better have learned how basketballs dribble, how light refracts through a scene, and all these little components of physics in order to paint this broader picture of the entire scene. So it's basically an emergent property from lots of compute and lots of data. That's a very banal thing to say these days. Everyone gets it. Better lessons work. But fundamentally, if you don't have these internal models of how the world functions, you will always be a worse model and thus have a higher loss than one that does. That is how we think about the optimization pressure which makes physics emergent from large scale video pre-training.

视频模型有价值数据 Valuable data for video models

Host

在语言模型世界,我们从整个互联网发展到如今非常有价值的数据,比如博士们做不同的事情。在视频世界中,有没有类似的东西,比如一个非常复杂的物理片段,对模型来说很棒?

In the LLM world, we went from just the entire internet to now incredibly valuable data like PhDs doing different things. Is there an equivalent in the video world of like a really complex piece of physics that is great for the model to see?

Bill Peebles

没错。我是说,如果你随便找一个视频研究员,给他们看一堆视频,问他们最喜欢哪个,很可能如果有一个体操套路,他们会说,“就那个。”但我认为这是一个非常有趣的问题。这有点难以回答,因为我觉得智能在视频中的表现方式与在语言中非常不同。例如,你可以想象一个讲座视频,有人在教微积分课程,那显然具有文本风格的智能,教你关于世界、数学、物理的深层概念。另一方面,当你想到这些体操套路时,除了做体操套路所涉及的规划之外,没有智力上的智能。但其中有太多有趣的细节,比如套路中发生的小碰撞,还有背景中所有人之间的移动。构建一个模型,不仅能模拟体操运动员,还能模拟整个场景中的每一个人。

Right. I mean, if you take any video researcher and show them a bunch of videos and ask them which is their favorite, chances are if there's a gymnastics routine going on, they're going to be like, 'That's the one.' But I think it's a really interesting question. It's somewhat hard to answer because the way intelligence manifests in videos, I feel, is very different than the way it manifests in language. You could for example imagine a lecture video where someone is giving a course on calculus, and that clearly has an intelligence which is of the flavor of text, teaching you about some deep concept about the world, mathematics, physics. On the other hand, when you think about these gymnastics routines, there's no intellectual intelligence there beyond the planning involved in doing a gymnastics routine. But there is so much interesting detail about the small collisions that happen as part of that routine, again about all of the people in the background shuffling amongst themselves. Building a model that can actually simulate not just the gymnast but every single person in the broader scene.

视频模型数据与多模态智能 Data for video models and multimodal intelligence

Bill Peebles

我们仍在摸索什么样的数据才能真正造就一个出色的视频模型,但这里面有很多多模态的成分,对吧?多模态指的是视频中存在各种不同的智能片段,而这些在文本等其他模态中不一定存在。

And so, you know, we're still trying to figure out what kind of data really makes for an amazing video model, but there's a lot to it that's very multimodal in a way, right? Multimodal in the sense of all these different pockets of intelligence which exist in video, but not necessarily in other modalities like text.

快问快答:去年观点转变 Quickfire round: changed minds in the last year

Host

我总喜欢在访谈结束时安排一个标准的快问快答环节,把一些过于宽泛的问题塞到最后。那么首先,我想知道你们每个人在过去一年里对 AI 领域的哪件事改变了看法。

Well, I always like to end our interviews with a standard quickfire round where I stuff some overly broad questions into the very end. So maybe to start, I'm curious for each of you, one thing you've changed your mind on in the AI world in the last year.

Bill Peebles

我认为我对某些事情的时间线加速了,而对另一些事情则推迟了。我觉得我们常常高估了消费者和采用率,以及人们学习这项技术的方式。我们在实际科学上可能远远领先于大众,但在构建可接受、可访问的产品界面以及如何向世界推广方面却远远落后。然后当你进入实际的企业用例时,还需要应对监管问题,才能让人们真正采用这项技术。所以作为一个产品人员,即使是在消费产品上,看到幕后为了让事情成为现实而付出的巨大努力,真的有很多工作要做才能让人们真正接受这些东西。

I think my timelines for some things have accelerated and for other things have delayed. I think we often overestimate consumers and adoption and how people learn this technology. We could be way ahead of the general world in terms of the actual science, but way behind in terms of building a product interface that's acceptable and accessible, and how we teach the world about it. And then when you get into actual enterprise use cases, there are regulatory things you have to battle through to get people to actually adopt this technology. So all of that, working as a product person, even on the consumer product, just seeing the crazy amount of work that happens behind the scenes to make things possible. There's just a lot to do for people to actually adopt these things.

Host

我在 AI 方面更新最多的观点是,如果一年前你问我,我在意的是看一部大片,还是看一部由 Sora 3 级别的模型生成的电影,我会说 Sora 3 的电影和那部大片一样好。我不需要任何背景信息。只要它理解了电影内容,有有趣的剧情,那就足够了。但有趣的是,我觉得我与这些模型互动得越多,当背后没有明确的创作意图时,我确实会感到一种空洞。这是我事先没有预料到的,部分原因是我喜欢制作这些模型。所以我有这种想法感觉很奇怪。但我真的被震撼到了,生成的内容感觉多么有趣和引人入胜,即使是 Cameo 也是如此,对吧?看到你认识的人,感觉棒极了。看到你的宠物 Rocket,我的狗,从昨天起已经在 Sora 里可用了。但即使知道某人只是有创作意图去拒绝一个样本,或者真正迭代一个提示和剧情,这些在底层生成中与他们的情感有更深层次的联系,你实际上确实能感受到这一点,这让我非常惊讶。我不确定是什么触发了我的这种改变。我曾经会完全相信我们机器人制作的纯 AI 生成内容,只要机器人足够好。但我现在真的认为即使是我也不会觉得那有吸引力,这对我来说很有趣。我不确定这到底是为什么。但总有一些人性的碎片需要通过这些生成内容来传达,才能使内容在某种程度上真正有意义。所以我在这方面确实更新了看法。

The aspect of AI I've updated most on is if you asked me a year ago, would I care if I watch some blockbuster movie versus some movie generated by a model that's Sora 3 caliber or something. I would say the Sora 3 movie would be just as good as the blockbuster one. I don't need any information about it. Assuming it's grokked cinematic content and interesting plots, that's sufficient. Interestingly though, I feel like the more I've interacted with these models, I actually do feel a hollowness when there's not clear creative intent behind it. Which is something I would not have expected to have felt a priori, in part because I enjoy making the models. So it feels weird for me to have that opinion. But I really have been struck by how much more interesting and compelling generations feel, even with Cameo, of course, right? Seeing the people you know and them feels amazing. Seeing your pet Rocket, my dog who's now as of yesterday available in Sora. But even knowing someone just had the creative intent to reject a sample, or really iterate on a prompt and the plot that they're actually using in the underlying generation has some kind of deeper connection to them emotionally, you actually do kind of feel that in a way that is very surprising to me. And I'm not sure exactly what triggered this change in me. I would have been very convinced by completely AI generated feeds that our bot made as long as the bot was good. But I now really do not think even I would find that compelling, which is quite interesting to me. I'm not sure exactly why this is the case. But there's still some morsel of humanity which is very important to communicate through these generations for content to actually feel meaningful in some way. So I've really updated on that.

基于Sora API构建什么 What to build on top of Sora API

Host

如果你们明天都离开 OpenAI,必须在 Sora API 之上构建东西,你们会去构建什么?

If you guys all left OpenAI tomorrow and had to build things on top of the Sora API, what would you go build?

Bill Peebles

游戏。

Game.

Host

比如一个让你置身于电子游戏中的游戏,你……

Like a game where you put yourself in a video game where you...

Bill Peebles

是的,抱歉。我本可以给出更多细节,但就像你明天就要走一样?

Yeah, I'm sorry. I could have provided more details than that, but like are you leaving tomorrow?

Host

这是我深思熟虑的。

Here's my very well thought out.

Bill Peebles

你为什么这么问?

Why do you ask?

Bill Peebles

是的,不,我认为有……我是在游戏中长大的。我是在游戏里学会编程的。我无法想象那个领域有多少疯狂的机会。很多在我制作游戏时非常困难的事情,比如必须寻找美术资源或自己制作美术,那是一个劳动密集型的过程,学习所有那些工具。我不认为这正是需要实现的,但什么样的游戏模型能通过这种生成技术实现,就像我们讨论 Cameo 那样,这是这里独特实现的东西。外面有什么样的东西?所以我对此感到兴奋,期待看到人们在这个领域能想出什么。

Yeah, no, I think there is... I grew up in gaming. I learned to program in gaming. I can't imagine how much crazy opportunity there is in that space. The idea that a lot of the things that were very difficult for me when I was building games, like having to source art or build the art myself, and that was a labor intensive process, learn the tools to do all that sort of stuff. I don't think this is exactly what needs to manifest, but what kind of gaming models are enabled by this generative technology in the same way that we're talking about Cameo is something uniquely enabled here. What kind of things are out there? So I'm excited about that and seeing what people can come up with in that space.

Host

完全同意。我的回答类似,但从不同角度出发,我认为这项技术最有趣的地方在于它将如何转变并创造新的媒介,而不仅仅是 AI 电影。我觉得那是最无趣的概念。有人在电影方面做很酷的事情,但全新的事物最有趣。所以我认为交互式,从 Thomas 的角度可能是游戏,从我的角度,也许是某种交互式叙事。还没有人真正做好这一点。有选择你自己的冒险书籍。我觉得那些相当流行。

Totally. My answer is similar from a different angle, which is that I think the most interesting thing about this technology is how it will transform and create new mediums rather than just an AI film. I think it's the least interesting concept to me. There are people doing cool things with film, but completely new things are most interesting. So I think interactive, which is maybe gaming from Thomas' point of view, from my point of view, maybe it's interactive storytelling of some sort. And no one's really nailed this. There are choose your own adventure books. I think those were pretty popular.

Bill Peebles

还有一些像 Netflix 上的《黑镜:潘达斯奈基》那样的电影,你有点……

And like some movies like Bandersnatch on Netflix where you kind of like...

Host

把这个想法推得更远,比如一些协作式的创意传说之类的东西。为什么你认为没有更多消费产品,比如有趣的消费产品,建立在这些模型之上?仅仅是因为它们直到现在才足够好吗?

Pushing that idea a lot farther, like some collaborative piece of creative lore or something like that. Why don't you think there have been more consumer products, like fun consumer products, built on top of these models? Is it just that they weren't good enough until now?

Bill Peebles

我认为构建消费产品真的很难。一个好的消费产品是最难的事情之一。

I think it's really hard to build a consumer product. A good consumer product is one of the hardest things.

Host

再加上技术如此新颖。而且不清楚什么才能真正流行起来。我认为这仍然是一个很难把握的组合。

Combined with the technology being so new. And it's unclear what actually sticks. I think that's still quite a hard combination of things to nail.

Bill Peebles

是的。

Yeah.

Host

我的意思是你们刚刚在顶层做了两次,所以我想……

I mean you've just done it twice at the top so I guess...

Bill Peebles

仍然困难。

Still difficult.

Host

那你会基于 API 构建什么呢?

What about something you'd build on top of the API?

Bill Peebles

哦,天哪。

Oh, man.

Host

我知道我们不能把你从模型中拉走,但如果必须的话。

I know we can't take you away from the models, but if we had to.

Bill Peebles

是的,我会去训练新模型,你知道的。可能更偏向科学方面或面向机器人。我认为那里有很多很酷的新工作要做。

Yeah, I would just go train new models, you know. Something probably more on the science side or robotics facing. I think there's a lot of cool new work to be done there.

Host

好吧,我想把最后一句话留给你们。

Well, I want to leave the last word to you guys.

结束语与推广 Closing Remarks and Promotion

Host

我通常会说大家可以去哪里了解更多,但如果我们的听众没用过 Sora,我会很惊讶。所以,也许我不知道是否有具体的研究想推荐给大家,或者很酷的新产品功能。话筒交给你。你想把听众引向哪里。请开始吧。

I usually say where can folks go to learn more, but I'd be genuinely shocked if any of our listeners didn't use Sora. So maybe I don't know if there's specific research you want to point people to or like cool new product features. The mic is yours. Anywhere you want to point our listeners. Please go ahead.

Bill Peebles

我们昨天刚推出了角色客串功能。所以,如果你用 Sora,你可以上传自己的形象,和朋友们一起创作视频。现在,你可以上传任何东西,比如宠物、无生命物体,你可以从 Sora 创建自己的 IP,并从中塑造一个角色。我有一个正在走红的泡菜角色,叫 Pickleton。所以,去客串 Pickleton 吧。就是这样。这是我的作品。

We just launched character cameos yesterday. So, you know, if you use Sora, you can upload yourself and create videos with your friends. Now, you can upload anything, you know, pets, inanimate objects, you can create your own IP from Sora and create a character out of that. I've got a Pickle that's going viral, Pickleton. So, go cameo Pickleton. There you go. It's my piece.

Host

我有个很疯狂的想法。好吧。我只是在想点什么。

I have a wild one. Okay. Just trying to think of something.

Bill Peebles

这是开放麦,请讲。

It's an open mic, please.

Host

开放麦是……在香港迪士尼乐园,有一个……

Open mic is... in Disneyland Hong Kong, there is...

Bill Peebles

好的开始。好的开始。

Good start. Good start.

Host

有一个……是幽灵公馆,但有点像……我记得叫魔法公馆。我不确定。他们稍微改编了一下。有一盏灯……这是一个黑暗骑乘项目。所以你穿行其中,有一盏灯……整个概念是这只猴子有一盏灯,当它倒出东西时,物体就活过来了。所以它倒出来,然后墙上会有一套盔甲,手臂会停下来跟你说话。

There is a... It's the Haunted Mansion, but it's like... I think it's called Magic Mansion. I don't know. They have a slight riff on it. And there's this lamp... It's a dark ride. So you just go through it and there's this lamp... The whole concept is this monkey has this lamp and when it pours out, the objects come to life. So it pours out and then there'll be like an armor on the wall and the arm will stop moving and talking to you.

Bill Peebles

我们其实刚刚就构建了类似的功能,你只需录制一个物体,然后说让它活过来。所以我觉得那个骑乘项目很酷,但回到角色客串,它确实有那种神奇的特性。

We actually kind of just built that, which like you just record an object and then you say come make it alive. And so I think check out that ride is very cool, but also back to character cameos, it's like actually it has that magical property.

Host

是的。我被说服了。

Yeah. I'm sold.

Bill Peebles

是的,我们暂时取消了邀请码限制。和朋友们一起来吧。Sora 和一群人一起玩更有趣。而且,去试试吧。给我们反馈。

Yeah, we've temporarily made it so you don't need an invite code to get in. Come in with your friends. Sora is way more fun if you're in there with a group. And, yeah, check it out. Give us feedback.

Host

是的。好了,各位,非常感谢。这非常有趣。我想确保你们回去继续构建 Sora,因为世界会从中受益更多。但说真的,这非常有趣。

Yeah. Well, guys, thank you so much. This is a ton of fun. I want to make sure I get you guys back to actually building Sora, because the world will far more benefit from that. But, seriously, this was a ton of fun.

Bill Peebles

非常感谢你们的邀请。香港迪士尼乐园见。

Thank you so much for having us. See you in Disneyland, Hong Kong.

Host

是的,说真的。

Yeah, seriously.

互动版:逐字朗读 + 针对本期提问 →