From Plausible to Precise: The Next Frontier in Image Generation
打开互动全文版(中英对照 + 朗读 + 问答)→高通研究员讨论文本到图像生成中剩余的挑战,包括可控性、质量和效率。
A Qualcomm researcher discusses the remaining challenges in text-to-image generation, including controllability, quality, and efficiency.
曾经是计算机视觉领域的前沿研究问题,如今文本到图像生成已经发展到几乎任何人都可以要求 AI 系统生成一张图片,而且效果看起来相当不错。但看起来不错与正确并非一回事。要求生成几个不同的人,模型可能会生成同一张脸的变体。要求特定的构图、身份或主体数量,它可能会忽略这些细节。向更高分辨率或本地生成推进时,质量、速度和内存很快成为约束。图像生成的下一前沿是缩小看似合理图像与精确结果之间的差距,使这些系统可控、高效、可靠,足以持续生成你真正想要的高质量图像。站在这一研究前沿的一位研究者是 Fati Periqi,高通技术副总裁,他的团队在今年的 CVPR(计算机视觉与模式识别会议)上发表了 20 多篇论文。以下是 Fati 解释为什么图像生成仍有许多难题需要解决。
Once a frontier research problem in computer vision, text to image generation has reached a point where almost anyone can ask an AI system for a picture and get something that looks remarkably good. But looking good and being correct are not the same thing. Ask for several different people and the model may generate variations of the same face. Ask for a specific composition, identity or number of subjects, and it may ignore those details. Push towards higher resolution or local generation, and quality, speed, and memory quickly become constraints. The next frontier in image generation is closing this gap between plausible images and precise results. It's making these systems controllable, efficient, and reliable enough to consistently produce a high-quality rendition of the image you actually ask for. One researcher at the forefront of this work is Fati Periqi, vice president of technology at Qualcomm, whose team presented more than 20 papers at this year's CVPR, the computer vision and pattern recognition conference. Here's Fati explaining why image generation still has plenty of hard problems to solve.
也许我们是在要求一个模型同时解决太多难题。想想当你生成一个包含几个人的场景时会发生什么。模型必须理解我要求建模的问题,然后决定应该出现多少人,确定他们应该放在哪里,比如场景的构图,推理他们之间的互动,因为如果有一个人,还有另一个人,很可能存在某种联系,保留我们给出的人的身份,比如这是我的女儿,这是我的儿子,我希望他们出现在画面中,而不是随便什么人,最后在单一过程中将所有内容渲染出来。所以也许这要求太高了。
Maybe we are also asking a single model to solve too many difficult problems at once. Think about what happens when you generate a scene with several people. Now the model has to understand the problem that I'm asking to model and then decide how many people should appear, determine where they should be placed like the composition of the scene, reason about their interactions because if there's a person, if there's another person most likely there is some connection, preserve the identity of the person we can give, okay this is my daughter this is my son and I want them to be in the picture not like any random person, and finally render everything together in a single process. So maybe this is too much.
我是 Sam Sharington,这里是 TwiML AI 播客。十多年来,我一直在通过这样的对话探索塑造 AI 未来的想法和创新,帮助你了解什么是真实的、什么是下一步、什么是重要的。让我们开始吧。
I'm Sam Sharington and this is the TwiML AI podcast. For over a decade, I've been exploring the ideas and innovation shaping the future of AI through conversations like this one that help you understand what's real, what's next, and what matters. Let's jump in.
是的,我认为这甚至是我看到的这次对话与去年对话之间的一个巨大差异。文本到图像模型的性能和能力有了显著提升。不仅是性能和能力,还有可及性,比如现在让你最喜欢的 LLM 生成一张图片,它会做得非常非常好。这确实引出了计算机视觉社区的一个问题:如果这个问题已经解决到这种程度,我们还有什么可做的?
Yeah, that's a big difference that I see between I think even this conversation and the conversation we had last year. The performance and capability of text to image models has improved significantly. And not just performance and capability, but I think accessibility like now ask whatever your favorite LLM is to generate an image and it will do a really really good job. And it does kind of beg this question of the computer vision community like what's what's left to do like if this problem is is you know solved to this degree what what's left for us to work on
这是个合理的问题。T2I 模型,也就是文本到图像生成模型或图像到图像生成模型,已经变得非常擅长生成非常逼真的图像。光照看起来自然,细节看起来正确,整体质量可能令人惊叹。就像我们做的许多事情一样,最初有兴奋,有出色的工作出现,但如果你深入挖掘,你会发现还有很多事情需要完成。一个是可控性。在同一图像中生成多个人时,人们看起来几乎相同,面孔有点融合在一起。此外,我们过去已经证明可以在用户设备上运行这类模型,你不需要依赖云服务提供商。我认为大多数模型仍然局限于 1K×1K 分辨率,但你想超越这一点,这是一个以前没有真正解决的挑战。
That's a fair question. T2I models, text image generation models or image to image generation models have become incredibly good at producing very realistic images. The lighting looks natural, the details look right, and the overall quality can be amazing. As with many things we do, there's initial excitement, there's great work coming up, but still if you look into that one dive deeper, you realize there are many things to still to be accomplished. One was controllability. Generating multiple people in the same image, people look almost identical, faces kind of blend together. And also we showed in the past we can run such models on users' devices, you do not need to rely on a cloud service provider. I think most of the models still are limited to 1K by 1K resolution, but then you want to go beyond that, that is a challenge which has not been actually addressed before.
好的,我从中听到的是我们取得了很大进展,但仍有工作要做。当你考虑这些工作时,一些大的类别包括可控性,即真正让模型专注于你描述任务的方式或专注于你想要的输出的能力。然后你提到了质量,即更少的伪影、更清晰的图像,然后你提到了效率。我们需要保持在我们设备上运行最新最强大模型的能力。所以这些听起来像是研究人员继续工作的三个大类别。
Okay, so what I'm hearing in there is that we've made a lot of progress but there's still work to be done. And when you think about that work, some of the big buckets include controllability, the ability to really get the models to focus on the way you describe the task or focus on the output that you want. And then you mentioned in there quality, so fewer artifacts, sharper images, and then you mentioned efficiency. We need to keep up with our ability to run the latest and greatest models on the device. So those sound like three chunky buckets for researchers to continue to work in.
是的,Sam,这是一个非常好的描述。谢谢。你说得很好。不仅是高通,社区也在努力确保这类模型,生成模型,能够匹配指导、质量、可控性期望或效率期望目标,就像真正的相机一样,但你是用你的语言、你的提示词来拍照。所以还有很多事情要做,这就是为什么我们今年在 CVPR 上发表并展示了多篇解决这些挑战的论文。
Yeah, that's a very good depiction, Sam. Thank you. You know, you said it very well. It's not only Qualcomm but the community also trying to make sure such models, generative models, can also match the guidance or the quality or controllability expectations or efficiency expectation goals of, you know, like a real camera, but you are taking pictures using your words, your prompts. So there are still many things to be done, and that's why we publish and present many papers addressing such challenges at CVPR this year.
我们要讨论的第一篇论文实际上涉及你给出的这个例子,面部属性的多样性。这篇论文叫什么?
The first one we're going to talk about actually talks about this example you gave, the diversity of facial attributes. What's this paper called?
论文的名字是《Disco:解决文本到图像生成中的身份问题》。
The name of the paper is Disco: Resolving Identity Issues in Text-to-Image Generation.
我要向听众提几件事。第一,我们会在节目说明中提供所有这些论文的链接,另外我们会尽量把论文中的一些插图放到视频中。所以如果你不是在 YouTube 上观看或观看视频,请留意这些,因为所有这些论文都有非常好的说明性示例,有助于跟上对话。那么 Disco 和 Artican,相对于我们讨论的这些类别,它们真正针对的是哪些?
I'll mention to folks listening that a couple things. One, we'll have links to all these papers in the show notes, but also we will, I'll try to get some of the illustrations from the papers into the video. So if you're not watching on YouTube or watching video, look for that because all of these papers have really good illustrative examples that will help with following the conversation. And so Disco and Artican, when you think about them relative to these buckets that we've talked about, which of these buckets are they really going after?
它们涉及可控性,以及如何训练 T2I 模型更好地与用户的指导对齐。让我更深入地谈谈 Disco,如果可以的话,Sam。当我们观察 T2I 模型,包括我们自己的模型时,有趣的是这不是一个图像质量问题。我之前提到的问题,比如我们要求模型生成面孔和一定数量的面孔,它却一遍又一遍地生成相同的面孔,几乎相同的面孔。所以从质量上讲,从图像质量上讲,如果你看像素、噪声等一切,它看起来是逼真的。但缺失的部分是,那些基础模型,那些出色的模型,并没有真正学会创建真正不同的身份,因为现有的训练目标主要侧重于真实感和匹配用户提示,但它们没有明确鼓励人与人之间的多样性,而这非常重要。所以这一观察引出了我们一个简单的问题:如果身份或面部外观或任何多样性本身成为一个优化目标会怎样?这引出了 Disco 背后的想法。我们没有创建全新的 T2I 模型,而是保留了底层模型,并用强化学习对其进行微调。我们设计了奖励,同时鼓励几件事。
They are about controllability and how to train T2I models better aligned with the guidance from the user. Let me go a little bit deeper about Disco, if it's okay, Sam. When we look at T2I models, at our own models also, what was interesting was that this is not an image quality problem. The problem that I mentioned before, like we are asking the model to generate faces and a certain number of faces, and it keeps generating the same faces, almost identical faces, over and over again. So quality-wise, image quality-wise, if you look at the pixels and noise and everything, it looks realistic. But the missing piece was that those base models, amazing models, had not really learned to create truly distinct identities because existing training objectives focus heavily on the realism and matching the user prompt, but they don't explicitly encourage diversity between people, and that is very important. So that observation led us to a simple question: what if the identity or facial appearance or any diversity itself becomes an optimization objective? So that led into the idea behind Disco. Instead of creating a completely new T2I model, we kept the underlying model and fine-tuned it with reinforcement learning. We designed rewards that encourage several things simultaneously.
比如,图像中的不同人物应该具有不同的身份,我们不希望出现重复的面孔。这就是我们在论文中提到的“图像内多样性”。而在同一模型多次运行、使用相似提示词时,我们也不应反复生成相同的面孔。这就是“运行间”或“图像间多样性”。这些是我们在训练或微调模型时明确加入的新目标。此外,我们还希望模型能生成正确数量的人物——如果我要求生成两个人,就应该是两个,而不是三个。我们也把这个目标纳入其中。当然,我们仍然保留了之前的图像质量目标。综合所有目标,我们使用了强化学习,具体来说是一种叫做“组相对策略优化”(GRPO)的方法。
For instance, different people within an image should have distinct identities. You don't want to duplicate faces. That's something we call in the paper as intra-image diversity. And across different runs of the same model with similar prompts, we should not keep generating the same faces. So this is inter-run or inter-image diversity. These are explicit new objectives when we train or fine-tune the model. Also, we want the model to generate the correct number of people—if I ask for two people, it should be two, not three. We incorporate that into the objective as well. Of course, we still have the previous image quality objective. Putting everything together, we use reinforcement learning, specifically something called Group Relative Policy Optimization (GRPO).
退一步说,我觉得面部多样性在生成多人图像时确实是一个重要目标,但它似乎只是众多可能影响图像的方式之一。为所有可能的方式调整目标似乎不仅困难,而且违背了“苦涩的教训”。为什么不直接收集更多包含大量人脸的数据,用更好的数据来训练模型呢?
Taking a step back, it strikes me that diversity of faces is an important objective if you're generating images with multiple faces, but it seems like one of many possible ways you might want to influence an image. Tuning the objective for all possible ways seems not just difficult but anti-bitter lesson. Why not just collect more data with lots of faces and train the model with better data?
从某种意义上说,我们正是这么做的,但我们证明了并不需要大量数据。你说得对,可能有很多属性:多样性、面部多样性是一个目标,人物数量是另一个。但假设我们想生成某个动作、姿态、位置或身体姿势,我们说的是,你可以在训练这类模型时,将这些目标纳入整体的感知或图像质量目标之上。但你不需要大量数据。这篇论文的知识点在于,你可以将所有内容整合到一个强化学习框架中,使得训练或微调这类模型变得可行且成本可控。顺便说一句,输入图像的质量也存在差异。此外,课程学习——从简单场景开始,逐步增加复杂度——能让强化学习更加稳定。
In a way, that is what we are doing, but we show that you do not need a lot of data. You are right—there might be many attributes: diversity, facial diversity is one objective, number of people is another. But then, say we want to generate a certain action, pose, location, or body pose of a person. We are saying that you can incorporate such objectives in addition to overall perception or image quality objectives when training such models. But you do not need a lot of data. The knowledge of the paper is that you can incorporate everything into a reinforcement learning framework, making it possible and affordable to train or fine-tune such models. By the way, there is a difference between the quality of input images. Also, curriculum learning—starting from simpler scenes and gradually increasing complexity—makes this reinforcement learning much more stable.
我记得多年前关于课程学习的讨论,那时它总是一个理论上的改进。如今看到它被纳入实际的训练算法中,令人兴奋。
I remember conversations I've had years ago about curriculum learning, and it was always a theoretical improvement. It's exciting to see it being incorporated into practical training algorithms nowadays.
是的,你说得完全正确。课程学习现在正在产生影响,不仅在我们的论文中,在其他几篇论文中我也看到了。我们正在讨论课程学习如何让多模态模型变得更好。从 Disco 论文中得出的一个更广泛的教训是,有时模型在训练中似乎只是错过了正确的目标,就像我提到的那些情况。这主要不是架构的限制,而是你如何向算法提供目标和训练数据。
Yes, you are absolutely right. Curriculum now is making an impact, not only for our papers but also in several other papers I see. We are talking about how curriculum learning makes multimodal models better. One broader lesson from the Disco paper is that sometimes the model seems to simply miss the right objective in training, like the things I mentioned. It is not mainly an architecture limitation, but it is how you provide the objective and training data to the algorithm.
我对你回答的总结是:是的,你可能想要控制的属性有很多,追求每个属性都有心智或人力成本。但从训练过程本身来看,它相当高效,与传统微调完全不同。它非常节省数据,而且你可以应用课程学习等技术来提高计算效率。
The way I would summarize your answer is that yes, there are lots of different attributes you might want to exert control over, and there is a mental or human cost to going after each of these attributes. But from the perspective of the training process itself, it's fairly efficient and nothing like traditional fine-tuning. It's very data-efficient, and you can apply techniques like curriculum to make it compute-efficient as well.
我想补充一点,因为这和我们现在讨论的内容相关。在正确的目标和正确的数据上进行优化绝对至关重要——这正是这篇论文的主题。但我们也可以想想,这些是生成模型。也许我们是在要求一个模型同时解决太多难题。
Something I'd like to add, because it's related to what we're talking about now. Optimization with the right objective and the right data is definitely critical—that's what this paper is about. But we might also think that these are generative models. Maybe we are asking a single model to solve too many difficult problems at once.
这其实是在问:我们能否优化所有我们关心的属性?这样做的话,我们是否在要求模型同时做太多事情?
That's kind of asking the question: could we possibly optimize for all the attributes we care about? In doing so, would we be asking the models to do too many things at once?
这是个好观点。也许我们不应该这样做。我可以给你一个例子。想想当你生成一个包含多个人的场景时会发生什么。回到那个核心例子,模型必须理解提示词,决定应该出现多少人,确定他们应该放在哪里(场景的构图),推理他们的互动——因为如果有一个人,很可能有另一个人与之有某种联系——并保持身份。我们可以说,好的,这是我的女儿,这是我的儿子,我希望他们出现在画面中,而不是随机的人。最后,在一个单一过程中渲染所有内容。也许这要求太高了。这就是我们在另一篇论文(R2-CAM 论文)中探讨的内容。与其让一个模型做所有事情,不如将规划与渲染分开,就像人类艺术家可能做的那样?这引导我们进入了这篇论文。
That's a good point. Maybe we shouldn't. I can give you an example. Think about what happens when you generate a scene with several people. Going back to that core example, the model has to understand the prompt, decide how many people should appear, determine where they should be placed (the composition of the scene), reason about their interactions—because if there's a person, there's likely another person with some connection—and preserve identity. We can say, okay, this is my daughter, this is my son, and I want them in the picture, not random people. Finally, render everything together in a single process. Maybe this is too much. That's what we explored in the other paper, the R2-CAM paper. Instead of asking one model to do everything, what if we separated planning from rendering, similar to what human artists might do? That led us into this paper.
在我们深入探讨 R2-CAM 之前,我想到了刚才关于 Disco 的评论,以及关于不同属性的想法。要记住的一点是,我们对模型路由器的思考方式已经成熟。也许你有一系列模型,分别针对你关心的不同属性,当你的提示词进来时,路由器会说:‘这个图像可能包含很多人,让我把它路由到针对这个场景调优的模型’,而不是试图让一个模型优化所有属性。这与你在 R2-CAM 中采取的方法不同,我们稍后会深入探讨,但这是 Disco 可以实际投入生产的一种方式。
Before we dive into R2-CAM, I'm thinking about that comment applied to Disco and the idea about different attributes. A thing to keep in mind is how we've matured the way we think about model routers. Maybe you have a suite of models for the different attributes you care about, and when your prompt comes in, you have a router that says, 'This image will probably have a lot of people, let me route it to this model tuned for that,' as opposed to trying to have a single model optimized for all attributes. It's a different approach to what you've taken with R2-CAM, and we'll dig into that, but it is a way that Disco could be put into production practically.
你提出了另一个很棒的观点,我们正在研究这个方向。这听起来更像是一个智能体编排的图像生成框架或流水线,对吧?取决于输入的提示词。
You brought up another amazing perspective, and we are working on it. This sounds more like an agentic orchestrated image generation framework or pipeline, right? Depending on the input prompt.
也许我们想生成一个写实场景或卡通场景,或者图像里可能有文字或人工生成的图表。所以对于所有这些,我们有不同的属性,比如多样性、面部身份。我们可能有专门的模型和这些模型内部的专门流程。所以如何挑选正确的那个,我们正在研究这样的智能体式流水线。根据输入,它会去确定正确的工具。这就像,如果你把 Disco 看作是多样性的一个实例,另一个实例是更好的文本生成。所以它可以根据你想应用的地方去找正确的那个,然后编排最终生成。我认为这将是终极解决方案。你知道,如果你真的想生成顶级惊艳的东西,一个尺寸可能不适合所有人。所以我们需要这样的专门化。
Maybe we want to generate, for instance, a realistic scene or a cartoony scene, or maybe there's text or some human-generated graph in the image. So for all of it, we have different attributes like diversity, facial identity. We may have specialized models and specialized processes within those models. So how to pull the right one is what we are working on, such as agentic pipelines. Depending on the input, it goes and determines the right tool. This is like, if you consider Disco as one instance of, let's say, diversity, another instance is better text generation. So it can go find the right one depending on where you want to apply them, and then orchestrate the final generation. I think this is going to be the ultimate solution. You know, one size might not fit everyone if you really want to generate something amazing, top of the line. So we need such specialization.
是的,山姆,这个观点非常好。所以,我再说一下,关于 Disco 的图像真的很有趣。我从来没想过——我不知道,也许我只是从未要求生成多人的图像——但确实有一些非常好的图像,你用一个提示词或几个提示词,让各种不同的模型,比如 GPT、Nano Banana、Flux、High Dream 等等,去生成多个人。你说得完全对,你看这些图像,就像同一个人,整个群体都是一个模子刻出来的。模型会这样,真是有点令人惊讶。当你谈到或思考评估,超越这种跨不同模型运行的想法时,关于这个模型的评估,有什么特别让你印象深刻的地方吗?
Yeah, that's a very good perspective, Sam. So, I will again, the images with regard to Disco are really interesting. Like I never really thought of—I don't know, maybe I've just never asked for an image with multiple people—but there are some really good images where you have a prompt or several prompts where you're asking a variety of different models, GPT, Nano Banana, Flux, High Dream, a bunch of them, to generate multiple people. And like you're absolutely right, you look at these images and it's like the same person, cookie cutter across the entire group. It's kind of surprising that the models do that. When you talk about or when you think about eval beyond this idea of like running across the different models, did anything in particular jump out at you in terms of evaluation for this model?
哦,是的。例如,为了这个目标,我们研究了现有的基准,然后我们整合了一个新的基准,专门帮助人们促进这里进一步的研究,并提供一些标准化。我们称之为“多样化人类”。论文中也有这个数据集和基准的链接,我们邀请大家去看看。所以我们创建了这个基准,并在该基准以及其他基准数据集上评估了 Disco。我们看到,当你明确施加这样的多样性目标时,分数,比如唯一面孔准确率、扣分分数,显著提高。Disco 大约 99,但我们开始用的那些没有明确多样性目标的模型非常低;可能有超过 10 到 20 个百分点的差距。所以我们在论文中讨论的这个解决方案,在文本到图像生成中增加了多样性的新分数,除了非常重要的人类偏好分数之外。评估输出质量也优于基础模型。
Oh yeah. For instance, for this goal, we looked into existing benchmarks and then we pulled together a new benchmark to specifically help people facilitate further research here and also provide some standardization. We call it Diverse Humans. It's also available in the paper; the link for that dataset and benchmark is there, and we invite everyone to take a look at it. So we created this benchmark and evaluated Disco on that benchmark and also on other benchmark datasets. We see that when you explicitly impose such a diversity objective, the score, for instance, unique face accuracy, deduction score, significantly improves. Disco is around 99, but the models we started with that don't have such an explicit diversity objective are very low; there's maybe more than 10-20 percentage point gap. So this solution that we also talk about in the paper adds a new score for diversity in text-to-image generation, in addition to human preference score, which is very important. Evaluating the quality of the output is also superior to the base model.
而且你构建奖励函数的方式很有趣,我想从评估的角度你也会做类似的事情。所以奖励的一个组成部分是你所说的图像内多样性,也就是说,看一张图像,面孔是相同的还是不同的?本质上,你拿你的图像,应用一个面部检测器,用边界框识别面孔的位置,然后提取这些面孔,嵌入它们,然后你可以对面孔进行成对相似度计算,得到从一个面孔到下一个面孔的距离,以确定它们是否都是相同的面孔,或者本质上得到面孔多样性的度量。
And it's interesting how you built the reward function, and I'm imagining that you would do a similar thing from an eval perspective. So one of the components of the reward is what you call intra-image diversity, and that is, you know, looking at one image, are the faces the same or are they different? Essentially, you take your image, apply a face detector that identifies where the faces are with bounding boxes, then extract those faces, embed them, and then you can do pairwise similarity across the faces and get the distance from one face to the next to determine if they're all the same faces or essentially to get a measure of the diversity of the faces.
完全正确。是的,我们在论文中讨论了这一点。甚至有一个漂亮的流程图。我邀请大家去看看。如你所说,应用了一个面部检测器,这是现成的。然后我们取这些检测到的面孔,将它们嵌入到某个规范空间中以便比较。然后我们计算成对相似度,这给我们该图像一个分数,即多样性分数。当然,那是一个分数,我们可以做同样的事情。假设我生成了一张图像,我有另一张相同提示词的图像,但初始化不同,种子点不同。提示词相同,我们运行它,然后我们得到一张不同的图像。
Exactly. Yeah, we talk about that in the paper. There's even a nice flow diagram. I invite everyone to take a look at it. As you said, there is a face detector applied, and this is off the shelf. Then we take these detected faces, we embed them into some canonical space to be able to compare them. Then we compute pairwise similarities, and that gives us a score for that image, which would be the diversity score. Of course, that is one score, and we can do the same thing. Let's say I generated one image and I have another image with the same prompt, but the initializations are different, different seed points. The prompt is the same, we run it, and then we end up with a different image.
所以你开始谈论下一篇论文,RT can,以及这个模型如何分解文本到图像生成的问题。给我们讲讲吧。
So you were starting to talk about the next paper, RT can, and essentially how this model kind of breaks down the problem of text-to-image generation. Talk us through that.
那么是什么促使了那篇论文?Disco 很棒。我们作为一个关键例子,当然讨论了这种身份多样性。但在另一篇论文中,我们指出,让一个模型试图做所有事情可能要求太高了。再次回到智能体式流程,也许像人类一样处理这些 AI 挑战更容易。我们不会试图自己解决所有事情。不仅仅是我们自己,即使在这个领域,艺术。一个艺术家处理这个问题不一定会逐像素思考。他们会想,你知道,主题是什么,背景是什么。
So what motivated that one? Disco is great. We, as a pivotal example, of course talk about this identity diversity. But then in the other paper, we are making the point that maybe it's too much for a model to try to do everything. Going back to agentic flow again, maybe it's easier to approach some of these AI challenges like a human being. We do not try to solve everything ourselves. Not just ourselves, but even in this domain, art. An artist approaching this problem wouldn't necessarily think about it pixel by pixel. They think about, you know, what's the subject, what's the background.
有一个规划,对吧?
There's a planning, right?
有一个规划方面。是的。是的。
There's a planning aspect to it. Yeah. Yeah.
在这篇论文中,我们基于这个想法。我们不是让模型做所有事情,而是将规划与渲染分开,类似于人类艺术家的做法。所以我们有两个组件。架构师,它不生成像素,而是为场景创建结构或构图。然后它决定,例如,如果图像中有一个人,它应该出现在哪里,如果有多个,他们应该如何排列,以便看起来自然和逼真。然后艺术家从那个构图结构开始,生成最终的逼真图像,当然,从一个目标来看,如果我们提供身份,也要保持身份。所以是的,这就像一个规划、一个架构师然后一个艺术家的框架。
In this paper, we build on that idea. Instead of asking the model to do everything, we separate planning from rendering, similar to how a human artist would work. So we have two components there. Architect, which doesn't generate pixels but instead creates the structure or composition for the scene. Then it decides where, for instance, if there's a person in this image, where it should appear, and if there are multiple people, how they should be arranged, so that it should look natural and realistic. Then Artist starts with that composition structure and generates the final photorealistic image, while of course, from one objective, also preserve identities if we provide identity. So yeah, this is like a planning, an Architect then an Artist type of framework.
底层技术方法也是像 Disco 一样使用强化学习。它也是基于 GPO 的方法。
The underlying technical approach is also using RL just like Disco. It's also GPO based approach.
你说得对。我们也有这个奖励函数,利用我们想要优化的不同目标。在这个目标中,我们有图像内和组内,比如图像间、人类感知分数和计数准确率。
You are right. We also have this reward function leveraging different objectives that we want to optimize. In this goal, we had intra-image and intra-group, like inter-image, human perception score, and count accuracy.
在 R2 中,我们也有构图目标,可以把正确的脸放到构图中的正确位置,这是 GRPO 奖励函数的一部分。还有脸的姿态,因为现在我们在构图——我的意思是,如果拍集体照,我们不希望一个人朝这边看,另一个人朝那边看。我们可以这么说。所以这些类型的东西都合并到奖励函数中,进入强化学习。然后我们优化它。模型现在学会如何同时拉取并优化所有这些。
Here in R2 we also have the composition objective, which can put the right face into the right place in the composition, and that's part of the reward function for GRPO. Also, pose of the face, because now we're composing—I mean, we don't want one person to look this way and another to look the other way if we're taking a group picture. We could kind of say that. So those types of things are all combined into the reward function that goes into reinforcement learning. Then we optimize that. The model now learns how to pull all of it and optimize all of it at the same time.
为了确保我理解正确,你有一个提示词。我看到的例子是“三个好朋友在火星上骑独角兽”。你把它传给所谓的“架构师”。架构师说:“好的,我想要三张脸,它们会在这张图片的这里、这里和这里,”然后大致规划出这个画布。然后这基本上成为你强化学习循环的基础。图像实际上都是生成的,但它们被优化以锚定到这个画布表示上。这样理解对吗?
So to make sure I understand this, you have a prompt. The one in the example I'm looking at is three best friends riding unicorns on Mars. You pass that to the what you call the architect. The architect says, "Okay, I want three faces and they're going to be here, here, and here in this image," and kind of plans out this canvas. And then that essentially becomes grounding for your reinforcement learning loop. The images are all actually generated, but they're kind of optimized to ground to this canvas representation. Is that the right way to think about it?
对。所以那个模型——架构师——知道如何放置这些脸,而艺术家——我们开始的渲染基础模型——现在经过微调,能够根据来自那些脸部区域质心的位置来遵循指导。但有两件事要记住。有一个离线的训练阶段。所以这个模型——架构师和艺术家——被优化得更好。但在推理时,首先我们取提示词,如果你想要,比如,一个特定人物身份出现在那个图像构图中,我们提供那些图像,架构师生成脸的位置。那里没有学习;它只是建立在之前学到的基础上。然后艺术家——微调后的模型——渲染它。所以这个 GRPO 是在离线微调领域,这肯定是一个重要的区别。
Right. So that model—the architect—knows where to position such faces, and the artist—the rendering base model where we started—is now fine-tuned to actually follow the guidance based on the locations coming from the centroid of those face areas. But there are two things to keep in mind. There is a training phase offline. So this model—the architect and the artist—are optimized to do better. But at inference time, first we take the prompt, and if you want, for instance, a special person identity to be in that image composition, we provide those images, and the architect generates the locations of the faces. There's no learning there; it just builds on what it learned before. Then the artist—the fine-tuned model—renders it. So this GRPO is on the fine-tuning offline field, so that's an important distinction for sure.
所以架构师的输出本质上是三个元组——三个 XY 点——在三个好朋友在火星上的情况下。
So the output of the architect is essentially three tuples—three XY points—in the case of three best friends on Mars.
三个人。是的。没错。
Three people. Yeah. Exactly.
对。然后那成为艺术家的输入,然后艺术家只做一次单次推理,或者,我猜它是一个扩散模型,所以它是迭代的。
Right. And then that becomes input to the artist, and then the artist just does a single-pass inference, or well, I guess it's a diffusion model, so it's iterative.
它会在内部迭代,但你说得对——在某种意义上它是一次单次传递,因为它不会去调用其他东西。它会这样做。
It's going to iterate internally, but you are right—it's like a single pass in the sense that it's not going to go call something else. It's going to do this.
是的,这很有趣。为什么有趣?我想有趣的是它居然有效,因为艺术家只得到位置。在训练阶段,艺术家得到的是真实的脸,对吧?不只是质心,还是说它在训练中只得到质心?
Yeah, it's interesting. Why is that interesting? I guess it's interesting that it works because the artist only gets like locations. It doesn't get—in the training phase, the artist gets actual faces, right? Not just centroids, or does it only get centroids in training?
另外,如果我们想说有一个特定人物,有代表特定人物的特殊标记。如果给出了,那么我们需要在训练阶段提供那些目标脸。但你知道,这样的脸可以被生成——我们也这么做了,也是通过另一个模型。这些不一定是真实的脸。为了自动化训练过程,训练非常高效。所以在训练时,那些脸被给出,提示词被给出,然后生成目标。然后我们评估是否有正确的数量,生成的脸的身份是否与初始身份匹配,以及其他目标,如计数准确性和姿态,因为姿态可能是那些提示词的一部分,还有颜色。所以在训练中,我们评估所有这些,把它们拉入这个奖励函数,然后通过 GRPO 优化奖励函数。一旦优化完成,现在艺术家知道如何根据这些输入位置生成更好的、渲染更好的图像。是的,架构师已经运行并提供那些位置。
Also, if we want to say that there's a special person, there are special tokens representing a kind of special person. If it is given, then we need to provide in the training phase those target faces. But you know, such faces can be generated—and that's what we did, also by another model. These are not necessarily real faces. To automate the training process, the training is very efficient. So in training time, those faces are given, the prompt is given, and then the target is generated. Then we evaluate whether there was the correct number, whether the identity of the generated faces matches with the initial identity, in addition to other objectives like count accuracy and pose, because pose could be a part of those prompts, and also color. So in training, we are evaluating all of them, pulling them into this reward function, and then the reward function through GRPO is optimized. Once this is optimized, now the artist knows how to generate better, render better images given these input sites. Yeah, the architect already ran and provided those locations.
对于这个评估,你有一堆提示词,然后是一堆与这些提示词对应的脸。你在图像中评估几件事。一是提供的脸的身份是否保留在输出图像中,还有描述性场景或提示动作是否反映在输出图像中。
And for this one in evaluation, you have a bunch of prompts and then a bunch of faces that go along with those prompts. And you're evaluating several things in the images. One are the identities of the provided faces retained in the output image, and also is whatever the descriptive scene or the prompt action reflected in the output image.
对,是的。
Right, yeah.
对,我们必须做到所有这些。我们想确保我们仍然与输入提示词对齐。例如,如果你说三个好朋友在火星上骑独角兽,就像论文中的例子,我们仍然看到三个人,有独角兽,而且像火星。所以这是与输入提示词的对齐;我们仍然强加这一点。我们还希望确保生成的图像是高质量的——这是人类感知分数——并且有三个人,不是两个朋友。所以配置——好的,我提到了三个朋友,但特定的朋友,不是随机的人。不是很多人骑任何十字架,但你知道,有点像给我另一个人。所以我们实际上提供了一些我们脸的例子,所以我们也希望我们的图片中也有我们的脸。所以这是身份脸匹配目标。还有另一个:我们希望他们可能看镜头,如果那是集体照。我的意思是,这再次是改进图像生成的一个例子。这并不意味着这是唯一的解决方案,或者你知道,welcome.com 正在展示什么是可能的。
Right, we have to do all of it. We want to make sure that we are still aligned with the input prompt. For instance, if you say three best friends riding unicorns on Mars, like the example in the paper, still we see three people and there are unicorns, and it's like Mars. So that is alignment to the input prompts; we still impose that. We also want to make sure that the generated image is high quality—this is human perception score—and there are three people, not like two friends. So configuration—okay, I mentioned three friends, but specific friends, not random people. Not many people riding any cross, but you know, kind of send me another person. So we actually provide some examples of our faces, so we also want our pictures to be also our faces to be in the generated picture. So that's the identity face matching objective. And also another one: we want them to maybe look at the camera if that's a group picture. I mean, this is again one example of how to improve the image generation. It doesn't mean that this is the only solution, or you know, welcome.com is showcasing what is possible.
身份保留是这种方法价值的重要组成部分,还是你认为它是一种独立于身份保留的有价值的方法?
Is identity retention a big part of the value of this approach, or do you see it as a valuable approach independent of identity retention?
哦,当然。再说一次,构图本身会看起来更自然,因为我们让架构师告诉我们图像的结构。绝对。但你知道,就像我之前说的,信息更大。信息是:嘿,社区,所有从事图像生成、文本到图像生成模型的人,看看如果你真的聪明地思考目标函数,比如 Disco 论文,并且简化任务,比如 R2 论文,你还能完成什么。不要试图一次做所有事情——比如我要生成构图和渲染——而是把这些过程分成更可管理的部分。
Oh, absolutely. Again, the composition itself would look much more natural because we ask the architect to tell us about the structure of the image. Definitely. But you know, like I said before, the message is bigger. The message is: hey community, everyone working on image generation, text-to-image generation models, look what you can also accomplish if you really think smartly about the objective function, like the Disco paper, and also simplify the task, like the R2 paper. Not trying to do everything at once—like I'm going to generate the composition and the render—but divide those processes into more manageable parts.
所以我们还有几篇论文专注于这种质量类别。Disco 和 Archon 也是可控性和质量。但我们接下来的两篇,Pixel Rush 和 Inverill,真正专注于输出质量。谈谈这两个在做什么。
So we've got a couple of papers as well focused on more of this kind of quality bucket. Disco and Archon are also both controllability and quality. But the next couple we've got here, Pixel Rush and Inverill, are really squarely focused on output quality. Talk a little bit about what these two are doing.
当然。
Absolutely.
前两篇论文,如你所说,是关于让图像生成更可控、融入不同目标,并让模型更容易做我们期望它做的事情。你提到的 Pixel Rush 和 Inver Field 这两篇论文,则转向了不同的目标。现在我们问的是,如何比之前更高效地生成图像。因为高通在之前的 CVPR 上提出了并展示了许多关于如何在手机上更高效运行这类模型的论文。但现在我们说,能否把它提升到下一个层次?因为之前我们一直在讨论,比如说 1K 分辨率,这也是目前甚至云端模型的限制。但问题是,我们能否做到 4 百万像素、16 百万像素的图像生成?挑战不仅在于模型运行的速度。顺便说一句,现有的解决方案可能需要 50 秒到几分钟不等,它们并不快。但还有内存挑战,因为当图像分辨率变大时,我们需要在设备上的某个地方保留扩散过程的潜在特征。这需要很大的内存占用。那么,我们如何运行一个模型,既能生成超大尺寸、高分辨率的图像,同时质量还得是真正的高分辨率——不只是采样图像的高分辨率,而是真实的细节、大量的细节——并且运行时间合理,不需要等 10 分钟(这是当前这类模型所需的时间),而且能在手持设备(如智能手机)的内存上运行。是的,这就是我们在 Pixel Rush 中的目标。另外,我们并不想显著改变现有的模型。这一点值得注意。有很多优秀的图像生成模型,其中一些来自大公司,我们都知道。我们不想让人们去微调这些模型。我们说,嘿,你仍然可以使用你拥有的任何这些模型,然后按照我们论文中讨论的流程,这样你就可以用那个模型生成,比如说,4 倍或 16 倍更多的像素。
The first two papers, like you said, were about making image generation more controllable, incorporating different objectives, and making it easy for the model to do the things we expect it to do. The papers you mentioned, Pixel Rush and Inver Field, shift gears into a different objective. Now we ask how to generate images much more efficiently than even what we did before. Because Qualcomm proposed and presented many papers at previous CVPRs on how to run such models much more efficiently on mobile phones. But now we are saying, can we push it to the next level? Because before, we've been talking about, let's say, 1K resolution, and that's the current sort of limit even for cloud models. But the question is, can we do 4-megapixel image generation, 16-megapixel image generation? The challenge is not only how fast you can run the model. By the way, such existing solutions may take anywhere from 50 seconds to minutes. They are not that fast. But there's also a memory challenge because when the image resolution gets larger, we need to retain the diffusion process's latent features in memory somewhere on the device. So that requires a big footprint. So how can we run a model such that we generate an extremely large, high-resolution image, and of course the quality has to still be real high resolution—not just a high-resolution sampled image, but real details, a lot of details—and it will run in a reasonable time, not like 10 minutes, which is the current time such model states require, and it would run on, let's say, the memory available on a handheld device like a smartphone. Yeah, that was our objective for Pixel Rush. Also, we didn't want to change the existing models significantly. That is something to note. There are wonderful image generation models, some of them by big companies, as we all know. We didn't want people to have to fine-tune those models. We are saying, hey, you can still use any of those models you have, and then follow our pipeline discussed in the paper, so you can use that model to generate, let's say, four times or 16 times more pixels.
那么请谈谈生成过程。你们的方法有什么不同?
So talk a little bit about the generation process. What's different about the way you've approached this?
是的,当然。所以它同样从提示词开始,有一个基础生成,就像任何模型,比如 FLUX 模型,然后生成一个基础图像。我所说的基础图像是指,比如 1K 图像。然后我们有这个级联上采样阶段。这是这篇论文的新颖之处。这就是为什么我说是 CVPR 论文。这个级联上采样接收这个图像,现在是 RGB 像素,不是潜在空间,然后它使用,比如说,任何现成的图像超分辨率解决方案。可以是双三次上采样或更智能的方法。它生成,比如说,更高分辨率的图像。所以当我们这样做时,我们现在有了,比如说,16 百万像素的图像,而不是 1 百万像素。我们有很多像素。然后我们取那个图像,应用一个编码器,一个 VAE,然后进入潜在空间。在那个潜在空间中,当然,如我提到的,我们关心内存。我们现在将那个潜在空间分成可管理的块;我们对它们进行分块。然后我们改进那些潜在空间特征。但之后,我们仍然在潜在空间中。我们通过 VAE 解码器进入像素空间。有一点我们需要非常小心:也有使用分块的方法,比如我要创建一块,然后另一块,再另一块。当你这样做时,你会产生伪影,可见的接缝。
Yeah, absolutely. So it again starts with a prompt, and there is this base generation like any model, it could be, let's say, a FLUX model, and then it generates, let's say, a base image. What I mean by base image is, let's say, a 1K image. Then we have this cascade upsampling stage. That is the part that is novel about this paper. That's why I say CVPR paper. This cascade upsample takes this image, which is now RGB pixels, not latent space, and then it uses, for instance, any off-the-shelf image super-resolution solution. It could be bicubic upsampling or something smarter. It generates, let's say, a higher resolution image. So when we do that, we now have, let's say, a 16-megapixel image, not 1 megapixel. We have a lot of pixels. Then we take that image and apply an encoder, a VAE, and then we go into a latent space. In that latent space, of course, as I mentioned, we are concerned about memory. We now divide that latent space into manageable chunks; we patchify them. So then we improve those latent space features. But then, we are still in the latent space. We go through a VAE decoder to the pixel space. Something we need to be very careful about: there are solutions also using patchification, like I'm going to take and create a patch, then another patch, then another patch. When you do that, you create artifacts, visible seams.
意思是当你在原始空间中这样做时,你会产生块。这里的不同之处在于你是在潜在空间中进行的。
Meaning when you do that in the original space, you create patches. What's different here is that you're doing it in the latent space.
完全正确。原因有很多。一是潜在空间在空间维度上比原始像素空间小得多。另一个是在潜在空间中我们可以引入噪声。这是利用噪声的聪明方式。是的,这是另一个原因。
Absolutely. There are many reasons. One is latent space is much smaller in spatial dimensionality than the original pixel space. The other one is in latent space we can induce noise. And that is a smart way of leveraging noise. Yeah, that is the other reason.
那么你在描述级联的图像中,有一个粗略潜在细化阶段,然后是高质量潜在。这两个是分开的潜在空间还是同一个潜在空间?这是什么意思?帮我理解一下。
And so you have in the image describing the cascade, a coarse latent refinement stage and then high-quality latent. Are these two separate latent spaces or is it one latent space? Like what does this mean? Help me understand, wrap my head around.
潜在空间是原始模型的潜在空间。它不是分开的空间。
The latent space is the latent space of the original model. It is not a separate space.
那么细化阶段做什么?
So what does the refinement stage do?
细化阶段添加一些引导性的语义噪声,然后迭代几次。这是新的部分,是我们提供的部分。所以它从这些潜在特征开始。通过注入的语义噪声获得额外的灵活性。我们想要这样,因为我们不希望,突然,好吧我们生成一个大图像,但图像中没有足够的纹理,或者高分辨率的语义上有意义的纹理。所以这就是为什么我们希望模型有灵活性来添加东西。我们在潜在空间中,然后我们引入这样的噪声。我的意思是,我们确实向这个潜在空间添加一些引导噪声,因为我们使用的扩散模型在潜在空间中进行扩散,它们从噪声开始,比如随机噪声,然后迭代很多很多次。我说的是扩散的高层概念。每次它估计一些噪声并将其从之前的噪声中移除,然后逐步澄清,最后得到一个无噪声的最终图像。它看起来像真实图像。这就是扩散去噪文本到图像模型的工作方式。所以我们在同一个空间中。潜在空间不是像素 RGB,而是特别低的,比如说输入是 1K 乘 1K,这可能是 128 乘 128。当然在每个点我们有一个向量表示对应像素的块的特征。那就是潜在。
The refinement stage adds some guided semantic noise and then iterates a couple of times. That is the new part, the part that we provide. So it starts with these latent features. It gets some additional flexibility through this injected semantic noise. We want that because we don't want, suddenly, okay we are generating a big image but then there isn't enough texture in the image, or high-resolution semantically meaningful texture. So that's why we want the model to have flexibility to add things. We are in the latent space and then we induce such noise. I mean, we add literally some guided noise into this latent space, because diffusion models, the ones we are using, they do diffusion in the latent space and they start with noise, like random noise, and then they iterate many, many times. I'm talking about the high-level idea of diffusion. Every time it estimates some noise and removes it from the previous noise, and then it clarifies step by step, and then it ends up with a final image which is now noise-free. It looks like a real image. That's how diffusion denoising text-to-image models work. So we are in the same space. The latent space is not the pixel RGB, but these are especially lower, let's say input is 1K by 1K, this is maybe 128 by 128. And of course at every point we have a vector representing the features for that patch corresponding to a pixel. So that is the latent.
潜在空间的维度是否比用于较小图像的更大,还是相同?
Is the dimensionality of the latent space larger than what you might use for smaller images, or is it the same?
是的,潜在空间维度是相同的。这就是为什么这个方法仍然非常高效。我的意思是,我们当然可以去更大的潜在空间,但那样会……
Yeah, the latent space dimension is the same. That's why this method is still very efficient. I mean, we could of course go to a larger latent space, but then...
但那是有代价的。
But that has a cost.
但内存挑战就在那里,对吧,而且也会花很长时间。
But the memory challenge is there, right, and also it will take forever.
所以从概念上讲,我对你们所做工作的理解是:你生成一张低分辨率图像,应用一个现成的上采样器得到更大的图像,但那张图像会有点模糊,质量也不太好,因为那是上采样的最先进水平。然后你基本上进入潜在空间,到达那个模糊的状态,再应用扩散来生成出更高清晰度的图像。
So conceptually, the way I'm thinking about what you're doing is you generate a low resolution image, you apply an off-the-shelf upsampler to get a larger image, but that image is going to be kind of blurrier and not all that great because that's the state-of-the-art for upsampling. And then you essentially go into latent space, and you end up in the blurry place, and then you apply diffusion to kind of generate out a higher definition.
当我们进入潜在空间时,有一个细微差别,而且很重要。我们对图像输入进行了上采样,这是 RGB,也就是彩色图像,对吧?现在我们有了一个更大的版本。它很大。如果我进入与目标尺寸成比例的潜在空间,那会非常慢。所以我们做的是将输入图像分割成原始大小的补丁。假设我有一个 1K×1K 的图像。现在每边有 4K×4K,所以我把它分成 16 个部分。于是我有 16 个 1K×1K 的图像,然后我……
When we go to latent space, there is a nuance and it's important. So we upsampled the image input, and this is RGB, like this is a color image, right? We have a larger version of it now. It is large. If I go to the latent space of the same proportion size to the target size, it's going to be very slow. So what we do is we partition the input image into original size patches. Let's say I have a 1K by 1K image. Now I have like 4K by 4K on each side, so I divide it into 16 parts. So I have 16 1K by 1K images, and then I...
你是把每个补丁投影到各自的潜在空间吗?
Do you project each patch into its own latent space?
当然。
Absolutely.
嗯。
Yeah.
明白了。
Got it.
新的那些。对。
The new ones. Yeah.
好的。那么你们在投影出来时如何避免边界处的伪影呢?
Okay. And then so how do you avoid artifacts at the borders when you project out?
所以这些图像并不是我们想要的那种超分辨率,因为它们只是上采样,对吧?我们实际上没有添加任何有语义意义的细节。顺便说一下,我们仍然有原始提示词。这就是为什么我们为每个潜在空间添加噪声。
So these images are not really super resolved in the way that we wanted because they are just upsamples, right? We really didn't add any semantically meaningful details. We still have the original prompt, by the way. And that's why we add noise for each latent space.
是的,这就是为什么我把它看作是对这个潜在空间应用扩散,你知道,文本驱动的扩散。
Yeah, that's why I think about it as like applying diffusion, you know, text-driven diffusion to this latent.
是的,是的。所以它允许我们创造更多的纹理。算法肯定会生成比输入图像更好看的输出图像,给定我们进入潜在空间之前的输入图像。所以在潜在空间中,我们现在为算法提供了这种自由,以生成更好的语义上有意义的纹理。但现在我有 16 个补丁,对吧?如果我把它们粘在一起,按原始排列排列,会有,因为它们彼此独立生成,对吧?由于内存限制,当然我们允许一些重叠,小重叠,但并不能解决问题,你知道,差不多。顺便说一下,另一个细微差别是那些补丁有轻微重叠,但并不能解决我们的问题。所以我们也意识到,当我们对这 16 个补丁的潜在空间数据表示进行混合时,我们喜欢在将它们组合在一起时再次进行另一次噪声注入。但在这种情况下,它不是像整个补丁一样,在潜在空间中相同,就像从同一分布中采样。我们有点允许算法在边界处生成更多噪声,但在补丁中心,你知道,也许程度较低。现在我们生成这个高质量的潜在表示,它同时解决了纹理问题,也解决了边界接缝问题,因为我们有……
Yes. Yes. And so it allows us to create more texture. The algorithm is going to definitely generate a much nicer looking output image than the input image, given the input image before we went to the latent space. So in the latent space now we've provided this freedom for the algorithm to generate better semantically meaningful texture. But now I have 16 patches, right? And if I stick them together, arrange them in the original arrangement, there will be, because they did it independently from each other, right? Because of the memory, and of course we allow some overlap, small overlap, but it's not going to solve it, you know, kind of. By the way, another nuance is those patches are overlapping slightly, but it's not going to solve our problem. So what we also realized is that when we are doing blending across such latent space data representation of these 16 patches, we like to again do another noise injection when we combine them together. But in this case, it is not like all over the patch, same in the latent space, kind of like sample from the same distribution. We kind of allow the algorithm to generate more noise towards the boundary, but in the center of the patches, you know, maybe to a lesser degree. Now we generate this high-quality latent, which at the same time now resolved texture but also resolves the boundary seam issue because we had...
所以你有点像,我在想象,不是独立地逐补丁处理,而是成对或交叉地细化这些补丁,以便在细化过程的另一端,它们在某种意义上彼此感知,边缘对齐并混合,诸如此类。
So you kind of like, I'm imagining that as opposed to independently processing patch by patch, you're like pair-wise or crosswise refining these patches so that on the other side of that refinement process, they're kind of aware of each other in a sense that the edges are aligned and blended and all of that kind of stuff.
是的。但当我们生成这个高质量的潜在表示时,大部分问题已经解决了。它将允许,当我们进入 VAE 解码器时,生成这个高分辨率图像,现在没有伪影了。我的意思是,我们解决了问题,在潜在空间中处理测试,那里更容易做,因为如果我们进入图像空间,如果我们开始在图像空间中混合,我们很可能会在接缝周围模糊,这是另一个我们不想要的伪影。
Yes. But when we generate this high-quality latent, most of the problem already solved. It will allow, when we go to the VAE decoder, generating this high-resolution image which is now artifact-free. I mean, we solve the problem, handle the test at the latent space where it is easier to do, because if we go to the image space, if we start blending in image space, we will most likely blur around the seams, which is another artifact we don't want.
嗯,这能行真是令人惊叹。
Well, it's amazing that that works.
哦,是的。我的意思是,再次强调,我们在这个案例中的所有动机是达到一个质量水平,就像我们有大的内存和大量的算力,我们可以在更大的潜在空间中做所有事情,但我们正在补丁化,并且可能做得更快。当我说快得多时,不是快两倍。可能是快 35 倍。你知道,从比如说 10 分钟到大约 20 秒那种加速。
Oh yeah. I mean, again, all our motivation in this case is to achieve a quality level like we have a big memory and a lot of compute power and we can do everything in a bigger latent space, but we are patchifying and do it maybe much faster. When I say much faster, it's not like two times faster. It's maybe 35 times faster. You know, from let's say 10 minutes to around 20 seconds type of acceleration.
有趣的是,我们在谈论边界和图像,以及正确重建它们,因为这也是逆变器在做的事情的一部分,不是吗?
And it's interesting that we're talking about kind of boundaries and images and like reconstructing them correctly, because that's part of what the inverter is doing also, isn't it?
在逆变场中,我们也有语义引导的噪声注入,就像像素冲刺一样。但场不是 T2I 模型。它是一个图像修复模型。所以这是图像到图像的生成式 AI。例如,你可以触摸一个物体,创建一个蒙版,然后移除那个物体,或者你可以将另一个物体,一个新物体带入场景,然后你可以创建这个非常实用的东西。因为很多时候我们拍照,但在背景中可能有一些我们想要移除的东西。挑战是,嗯,我仍然可以看到,你知道,你创建纹理有时还可以,它是有意义的,但然后……
In invert field, we also have a semantically steered noise injection like the pixel rush. But field is not a T2I model. It is an image inpainting model. So this is image-to-image generative AI. For instance, you can touch an object, create a mask, and then remove that object, or you can bring another object, a new object into the scene, and then you can create this very practical thing. Because many times we take pictures, but in the background maybe there are things that we want to remove. The challenge is, well, I can still see, you know, you create the texture sometimes okay, it is meaningful, but then...
你在背景中看到的那些伪影,比如……
There's those artifacts that you see in the background like the...
海滩上的沙子,你移除了不需要出现在照片中的人,看起来有点古怪。
Sand on the beach where you remove the person who didn't need to be in the picture is kind of funky looking.
而且也许纹理与图像的其余部分不太一致,或者我确实看到边界周围的伪影。所以现在,是的,它就像一个玩具,你知道,你可以那样做,但我不会真正使用它。但我们说的是,嘿,你不需要那样,你可以做更好的物体移除或图像修复、图像编辑,这就是这篇逆变场论文所讨论的。
And maybe the texture is not really compliant with the rest of the image, or I see literally the artifacts around the boundary. So car now, yeah, it is like a toy, you know, you can do that, but I'm not going to really use it. But we are saying that, hey, you don't need to be there, you can do much better object removal or image inpainting, image editing, and that's what this invert field paper is talking about.
那么这篇论文背后的核心思想是什么?我在想象你使用了扩散,你知道,也许在某个地方使用了多步或单步扩散。
So what's the core idea behind this paper? I'm imagining that you're using diffusion, you know, maybe multi-step or one-step diffusion somewhere.
所以在逆变场中,我们做的是,好的,我们有这个输入图像,我们允许在背景中生成噪声,然后我们也有这个在蒙版中的噪声,比如我们想要移除的鸟。但我所说的背景中的噪声是指图像的其余部分。所以在去噪过程中,我们在训练期间从带噪声的图像开始,然后逐步去除噪声,最终得到干净的图像。在图像生成中,我们也做同样的事情,对吧?我们从噪声开始,最终得到干净的图像。但想想另一个过程,就像我们实际训练那些模型的方式。那些模型,当我们训练时,我们从干净的图像开始,然后添加噪声,最后变成带噪声的图像。
So in invert field, what we do is, okay, we have this input image and we allow generating noise in the background, and then we have also this noise in the mask, like the bird that we want to remove. But what I mean by this noise in the background is the rest of the image. So in denoising, we started with, during training, with a noisy image, and then we progressively removed noise and ended up with a clean image. In image generation, we also do the same thing, right? We start with noise and then end up with a clean image. But think about the other process, like the way that we actually train those models. Those models, when we train, we start with the clean image and then add noise, and at the end it becomes like a noisy image.
所以想想我们实际训练模型的反向过程。我们可以把这张输入图像映射成噪声。我们是在逐步把一个真实的、比如说是干净的图像反转成带噪声的版本。这是研究得很透彻、理解得很清楚的东西,而且速度非常快——大约 60 毫秒——我们可以拿一张大图,然后通过反向去噪过程,最终得到一个它的带噪声版本。但这个噪声不再是随机噪声了,它是特定于输入图像的。所以如果我改变输入图像,噪声也会不同。这是针对整张图像的,然后我们有这个鸟的掩码。现在我对鸟的部分添加了噪声,因为我想让算法生成一只符合我文本提示的新鸟。所以现在我改变了我们创建这张图像的方式,但我不需要改变原始模型。我仍然可以改变我在这个绘画中初始化扩散的方式。当我们这样做时,当我们从这个反转噪声加上掩码以及掩码内的新噪声开始时,首先,我们可以保留背景,但我们也允许背景稍微影响前景,就像掩码本身一样。这实现了无缝的协调,并生成高质量的图像。但最重要的是,不再有边界伪影了,因为我们在同一图像中有两种噪声——随机噪声和反转噪声。当我们开始生成图像时,我们不是仅仅从掩码内的噪声开始的。
So think about the reverse process where we actually train the model. We can take this input image and map it into noise. So we are progressively inverting a real, let's say clean image into noisy versions. This is something very well studied and understood, and it's very fast—like 60 milliseconds—we can take a large image and then end up creating, going through the reverse denoising, a noisy version of it. But this noise is not random noise anymore; it is specific to the input image. So if I change the input image, the noise is going to be different. So this is about the entire image, and then we have this mask of the bird. Now I added noise to the bird because I want to allow the algorithm to generate a new bird compliant with my text prompt. So now I changed the way that we create this image, but I don't need to change the original model. I can still change the way that I initialize the diffusion in this painting. And when we do that, when we start with this inverted noise plus the mask and new noise within the mask, first of all, we can retain the background, but we also allow the background to slightly impact the foreground, like the mask itself. This allows seamless harmonization and generates high-quality images. But most importantly, there are no boundary artifacts anymore because we have two noises—random noise and inverted noise—within the same image. When we start the image, we are not starting from just noise within the mask.
太棒了。那么,正如我们一开始提到的,高通在 CVPR 上总是有很多论文。我们无法全部覆盖。我们已经介绍了几篇最重要的图像生成论文。但还有关于视频生成的论文。还有很多演示。有没有什么特别想让大家关注的其他内容?
Awesome. So, as we suggested starting up, Qualcomm always has a ton of papers at CVPR. We can't cover all of them. We've covered a handful of the most important image generation papers. But there were also papers on video generation. There were a ton of demos. Anything in particular you'd want to call out in terms of other things folks should look for?
当然。我为这三篇视频生成论文感到非常自豪,因为它们让视频生成对每个人都变得可及。你现在可以在你的 PC、笔记本电脑或手机上运行这样的模型。我只提一下它们的名字。第一篇是 Preal——它是最好的开源、公开可用的模型之一。所以我们在那篇论文中展示了你可以实际运行得更快,比如快五倍,然后把它压缩到内存有限的设备中。你不需要在云端运行它。另一篇论文是关于混合注意力——循环混合注意力。同样,这里存在在哪里应用注意力的挑战,所以它做得更好,也更快。还有注意力手术,因为注意力消耗大量算力。所以这些论文都是关于视频生成的。很棒的论文。请也看看那些项目页面。我们为这些论文准备了项目页面。
Absolutely. I'm very proud of the three video generation papers because they make video generation accessible to everyone. You can run such models now on your PC, laptop, or your phone. I will just mention their names. The first one is Preal—it's one of the best open-source, publicly available models. So we are showing in that paper that you can actually run it much faster, like five times faster, and then squeeze it into your memory-limited device. You don't need to run it on the cloud. The other paper is about hybrid attention—recurrent hybrid attention. Again, there's this challenge of where to apply attention, so it does a much better job and makes it faster. Also, attention surgery, because attention takes a lot of compute. So these papers are for video generation. Amazing papers. Please take a look at those project pages also. We have project pages for these papers.
好的,Fatih,一如既往,和你交流非常愉快,听你介绍你们在计算机视觉方面的工作以及你在 CVPR 上做的所有酷炫的事情。
Well, Fatih, as always, it's been great catching up with you and hearing about what you all are doing with regards to computer vision and all the cool things you did at CVPR.
Sam,谢谢你邀请我。能参加你这个精彩的播客总是非常荣幸。我很高兴能参与其中,谈论这些作品。我只提到了其中几个,我期待再次见面,再次参加你的精彩播客。非常感谢你的邀请。
Sam, thanks for having me. It's really always a pleasure to be a part of your amazing podcast. I'm very excited to be a part of it, talk about artwork. I only mentioned a few of them, and I look forward to meeting again and joining your amazing podcast. Thank you so much for inviting me.
好的。谢谢你。
All righty. Thank you.