The Infinite Canvas: A Pivotal Interface for AI Collaboration
打开互动全文版(中英对照 + 朗读 + 问答)→tldraw 创始人 Steve Ruiz 探讨为何无限画布是人机协作的强大基础,并分享他从美术到构建画布工具的历程。
Steve Ruiz, founder of tldraw, discusses why the infinite canvas is a powerful primitive for human-AI collaboration, sharing his journey from fine art to building a canvas that powers many startups.
我越是观察工具领域的格局,就越确信无限画布将在我们与 AI 的交互中扮演关键角色。如何实现人类 + AI + 多个 AI 与多个人类在单个文档中实时协作?所以,我想深入探讨为什么画布是如此强大的原语,以及设计优秀工具所伴随的所有隐藏复杂性。很少有公司、产品或人在解决这个问题——我们处于一片巨大的未知领域,但这也是软件领域最雄心勃勃的事情之一。欢迎来到 Dive Club。我是 Rid,这里是设计师永不停歇学习的地方。本周的嘉宾是 tldraw 的创始人 Steve Ruiz,他花了三年多时间和 500 万美元构建了完美的画布,这个画布驱动了我们节目中研究过的许多初创公司。所以,我们将深入探讨画布用户体验以及它为 AI 未来解锁的一切。但我最喜欢这次对话的部分之一是听到 Steve 的旅程,因为它始于一个你可能意想不到的地方。
The more that I look at the landscape for tooling, the more convinced I am that the infinite canvas will play a pivotal role in how we interface with AI. How do you do human plus AI plus multiple AIs and multiple humans collaborating in a single document in real time? So, I wanted to dig into why the canvas is such a powerful primitive and all of the hidden complexities that come with designing a great tool. There are very few companies, products, anyone working on that problem that we are in like massively uncharted territory, but also like some of the most ambitious stuff happening in software. Welcome to Dive Club. My name is Rid and this is where designers never stop learning. This week's episode is with Steve Ruiz, who is the founder of tldraw, where he spent over three years and $5 million building the perfect canvas that powers a lot of the startups that we've studied on this show. So, we're going to do a deep dive into Canvas UX and all of the things that it unlocks for the future of AI. But one of my favorite parts of this conversation is just hearing about Steve's journey because it starts in a place that you might not expect.
我本科和硕士都学的纯艺术。我仍然会画画。我对此非常感兴趣,不仅作为一种手艺,也作为一种职业和行业——视觉文化是如何产生的?艺术是如何创作的?艺术家的生活是怎样的?我如何参与其中?但我也如何支持它?在大学和研究生之间,我做了很多艺术写作,在芝加哥为不同的展览写艺术评论和新闻报道。我在芝加哥大学读了两年,那是一个非常概念性的项目,所以不仅仅是手艺,而是非常学术的。大学期间,我遇到了一个可爱的女人,她现在是我的妻子。她在剑桥找到了一份工作,所以我搬到了英国,我们结婚了,我有了一个工作室,在英国做艺术。你知道,创意职业非常脆弱。所以,在工作室待了大约 6 个月,或者更长一点,大约一年后,事情并没有按照我想要的方式发展。我甚至不确定我是否还希望它们按照我想要的方式发展。在芝加哥的时候,我一周部分时间给律师工作,另一部分时间在工作室。那算是一份高薪的日常工作,基本上是做法律研究,现在可能已经被 ChatGPT 取代了。你知道,“嘿 Steve,这里有一个法律问题,去法律图书馆查清楚。”英国有不同的法律,说实话,我在落地之前并没有想到这一点。我以为我能继续做这类事情,保持那种分析性的一面(每周三天)和创造性的一面(每周五天)。但我无法从事那种角色。所以我有很多多余的分析性心理能量,对吧?这两方面我一直保持平衡。所以我想,你知道吗?我不想回学校,不想重新培训去做一些我并不真正热衷的事情。我从未真正尝试过将自我的这些不同部分结合起来。我一直对找创意工作很抵触,因为我认为那是我在工作室里保留的东西,用于创作,这是真的。你每周只有那么多创意精力。所以,我想,也许是时候了,让我关闭工作室,试着找一份能让我利用这两方面的工作。我想在剑桥,最明显的方向是出版业和出版业内的设计。我学会了使用 Adobe InDesign,因为我知道那些出版社会教这个。我开始面试,最终得到了一份工作。我整天用 Adobe InDesign,心想:“好吧,不错。这是一个好的开始。我有了工作。”但让我继续深入设计。我想成为什么样的设计师?我的心思在出版业,因为那是剑桥存在的机会。我实际上对电子书非常着迷。电子书太棒了。它们就像 HTML 和 CSS,打开电子书,弄清楚它们是如何工作的,如何排版的,如何打包和分发的,等等。所以我非常喜欢电子书,然后我想,这基本上就是网页设计,对吧?我们用的是 HTML CSS。我在谈论移动东西。我有一点 HTML 技术,因为大学期间我为其他艺术家和自己制作过 WordPress 网站。所以我有一些技术知识。那是 2016 年。一些非常有趣的事情正在发生。设计师在网页设计中的角色已经出现,这在 10 年前并不存在。至少据我所知,它更加混合。还有一种产品设计的事情正在发生。真正优秀的设计师正在构建实际可用的东西。你知道,那时 Uber 和 Airbnb 开始利用原型工具,Origami 已经发布,Framer 非常早期的版本,你知道,带有 JavaScript 库后的 IDE。我想,哦,好吧,这才是真正专业化的东西,因为我知道我需要某种专业化,但我也觉得我不想以一种无聊的方式进入这个领域,我真的想深入其中,结果原型设计成了我的切入点。所以主要是通过 Framer Classic,或者后来成为 Framer Classic 的东西,现在它已经不存在了。所以是通过 Framer 的早期版本,那时它还不是一个网站设计工具。它就像一个带有代码的原型工具,左边是代码,右边是屏幕。它就像一个小型 IDE。我使用 CoffeeScript,它试图尽可能简单,让设计师说:“这是我的 Sketch 文件。让我标记我的图层,然后我可以说,好的,图层这个按钮。”
So I studied fine art at university undergrad and then again for my masters. I can still paint and draw. I mean it's something that I was really interested in not just as a craft but also a career and kind of an industry of like how does visual culture get made? How does art get made? What are the lives of artists like? How do I participate in that? But also how do I like support that kind of between university and grad school like I did a lot of art writing so I was like doing art reviews and journalism around different exhibitions at the time in Chicago. I did two years at University of Chicago which was very kind of conceptual program so it wasn't just about the craft it was really very very academic. During university I met a lovely woman who is now my wife. She got a job at Cambridge so I moved to the UK, we got married, and I got a studio and I was doing my art in England. You know, creative careers are really, really fragile. So, after maybe 6 months of being in my studio, or maybe a little longer than that, about a year of being in my studio, things aren't really moving the way that I want them to. And I'm not even sure if I wanted them to move the way that I wanted to anymore. When I was in Chicago, I was doing like basically working for lawyers for part of the week and then working my studio for the other part of the week. It was like a high-paying day job so to speak like to be doing legal research basically something that would probably be being replaced by ChatGPT right now. You know, hey Steve here's a legal question go to the law library and figure it out. England has different laws which is honestly was not something that I thought about before I landed. I was like oh yeah I'll just be able to keep doing like you know this type of thing and have this like kind of analytical side that I used for three days a week and then maybe like the creative side that I use for 5 days. I wasn't able to work in that type of role. So, I was just I had a lot of like excess analytical mental energy going on, right? Like those were always kind of two things that I had kept in balance and stuff. And so, I'm like, you know what? I don't want to go back to school. I don't want to like retrain to do something that I wasn't super super passionate about. Anyway, I'd never really tried to combine these different parts of myself. I'd always been very defensive about getting a creative job because I thought that that was something that I kept in my studio for work, which is true. You only have so much kind of creative juice during the week. So anyway, I was like, you know what, maybe it's time like, let me close my studio and let me try and find a job that lets me take advantage of both things, both parts of myself. I guess the obvious place for that when we were in Cambridge was publishing and design within publishing. Learned how to use Adobe InDesign because I knew that people were going to teach that at these publishing houses. Started going out for interviews and eventually got a interview, got a job now. I was working in Adobe InDesign all day and I was like, "All right, cool. This is a good start. You know, I'm employed." But let me keep going into design. What kind of designer do I want to be? My head was in publishing because that's like the kind of the opportunities that existed in Cambridge. And I got really into ebooks actually. Ebooks are fantastic. They're like HTML and CSS, like cracking open ebooks and like figuring out how they work and how they were typeset, how they were packaged and distributed and all that stuff. So, I got really into ebooks and then I was like, you know, this is all basically just web design anyway, right? Like we're HTML CSS. I'm talking about like moving things around. And I had a little bit of like technical HTML skill from just making like WordPress websites for other artists and myself like during college. So, I had like a little bit of that technical knowledge. This was like 2016 at the time. Some really interesting things were happening. The role of the designer within web design had emerged that really wasn't there like 10 years before. It was much more kind of hybrid at least as I understood it. And there was this kind of like product design thing that was happening. The really good designers were building stuff that actually worked. You know, this was when like Uber and Airbnb were starting to leverage prototyping tools like Origami was out, Framer very early Framer versions were you know with the IDE after the JavaScript library. And I was like oh okay this is what's really specialized because I knew that I needed some sort of specialization but I also was like I don't want to get into this in a way that is boring or in a way that is like I really want to get into this that turned out to be the kind of the my wedge or my way in was through prototyping. So it was mostly through Framer Classic or what became Framer Classic which now doesn't even exist. So it was through early versions of Framer before it was a website design tool. It was like a prototyping tool with code, like code on the left, like a screen on the right. It was like a little IDE. I was using CoffeeScript and it was trying to be as simple as possible for designers to say like here's my Sketch file. Let me label my layers and then now I can say like, okay, layer this button.
如果你早期就掌握了这些工具,那确实是一个差异化因素。而且当时你必须对复杂性有很强的接受能力才能学会,这很能说明你作为设计师的能动性和决心。影响力很大,我希望今天进入这个行业的人也能找到类似的路径。但这成了我的方向。我当时想,好吧,这就是我要变得非常擅长的事情,因为没人擅长这个。大家似乎都感兴趣,但没人能做。我那份排版工作,我并不喜欢,但合同还有 6 个月到期,我打算在 6 个月内找到一份全职工作,做我觉得有趣的事情。我大约 3 个月后开始接兼职,然后夏天找到了第一份全职工作,最终在 6 个月内进了一家普通的创业公司。
And if you got in early with those tools, it really was a differentiating factor. And also you had to have an appetite for complexity back then to be able to learn that, which I think says a lot about the level of agency and determination that you have as a designer. Highly influential and I hope that there's a parallel today for people coming into the industry. But that became my thing. I was like, all right, this is going to be the thing that I'm going to get really good at because no one's good at this. Everyone seems to be into it but no one can do it. This typesetting job that I have, I don't really like, but my contract is up in 6 months, let me try and find a full-time job in 6 months doing this thing that I find interesting. I started contracting after like 3 months on the side and then got my first full-time job over the summer and eventually got a normal startup job within 6 months.
快速插播一条消息,然后我们继续。我前几天看到一条让人停下滑动的推文。Tailwind 的创建者正直接与 Paper 合作,训练输出达到完美。他们甚至投资了这家公司。所以想想可能性。未来,你可以在 Paper 中设计东西,然后右键复制出完美的 Tailwind 代码,就像创建者亲手写的一样。或者你可以把现有的代码组件导入 Paper,直接在画布上编辑。这将彻底改变我们设计和交付 Web UI 的方式。这也是我大力押注 Paper 成为下一个伟大设计工具的又一个原因。你现在就可以试试。访问 dive.com/club/paper。如果你像我一样,就知道为设计添加动效是让它们显得高级的最简单方法。问题是我不是动效设计师。但这就是 Jitter 的新 AI 头脑风暴功能如此颠覆的原因。我只需放入设计,就能立刻获得动效创意,然后调整、优化,变成自己的。用 Jitter 给你的作品添加动画真的超级简单。我强烈推荐。访问 dive.com/club/jitter 试试吧。好了,现在回到正题。
Real quick message and then we can jump back into it. I saw a scroll-stopping tweet the other day. The creators of Tailwind are working directly on Paper to train the output to be perfect. They even invested in the company. So just think about the possibilities for a second. In the future, you could design something in Paper and then just right-click and copy the perfect Tailwind as if the creators themselves wrote it by hand. Or maybe you take an existing code component and import it into Paper to make edits directly on the canvas. And this is going to totally change how we design and deliver UIs for the web. And it's just another reason why I'm betting big on Paper as the next great design tool. You can try it out today. Just head to dive.com/club/paper. If you're like me, then you know adding motion to your designs is the easiest way to make them feel premium. The thing is, I'm not a motion designer. But that's why Jitter's new AI brainstorm feature is a game changer. I just drop in my design and then get instant motion ideas that I can tweak, refine, and make my own. It's seriously so easy to animate your work with Jitter. I cannot recommend it enough. Just head to dive.com/club/jitter to try it out today. Okay, now on to the episode.
我想把故事往前快进一点,因为你有艺术背景。你下定决心学习原型设计,以那种方式突破,而且没过多久就开始做画布 SDK。那么,你是怎么走到“我要创办一家画布公司”这一步的?
I want to zoom ahead of the story a little bit because you have this art background. You make the determined effort to learn prototyping, break in that way, and it didn't take you that long to start working on a canvas SDK. So, how the heck did you get to the point where you're like, I'm going to start a canvas company?
好吧。神奇的是,我的第一份技术工作是 2017 年。我在 2021 年开始做 Teal Draw,对吧?所以我 2017 年学会编程,之后四年里一直在做一些相当棘手的开发工具类工作。然后五年内就创办了公司。所以可以说我的职业生涯非常非常快,因为我一直处于那种原型设计师的世界,那就是我的角色。在代理公司工作时是这样,在产品公司工作时也是这样,最后我去了 Framer,为设计师做教育和内容,有点像试图培养出更多像我这样的人。
All right. So, that's the wild thing is that my first job in tech was 2017. I started working on Teal Draw in 2021, right? So I learned to code in 2017 and then I was working on some pretty gnarly dev tool type stuff within four years of that. Then I had the startup within five. So it was a really, really fast career, so to speak, because I was kind of in this world of being the designer who prototypes, like that was my role. It was my role when I worked in an agency, it was my role when I worked in product, and eventually I ended up working at Framer doing education and content for designers, like trying to make more Steves in a way.
在所有这些角色中,我都在非常非常积极地学习,同时也大量演讲、写作、教学、举办工作坊等等。所以我的知识增长得非常快,尤其是当我试图教别人的时候。我想,‘好吧,现在我必须真正了解 React 怎么工作。我必须知道这个那个怎么运作。’我认为如果你是一个初级开发者,你可能会被慢慢引入行业,别人会说,‘好的,这里有一个小任务,或者一些简单的事情你可以做。’你没有太多机会做真正有雄心的项目,因为我做的所有原型设计都是非常有雄心的短期项目,我能很快得到反馈,然后继续下一个。
Throughout all those roles, I was just really, really aggressively learning and also speaking a lot about it and writing a lot about it and teaching, running workshops and things like that. So my own knowledge just expanded really quickly, especially once I was trying to teach other people. I was like, 'Okay, now I got to really know how React works. I really have to know how this other thing works.' I think if you're a junior developer, you kind of get slow-balled into the industry a little bit by just saying, 'All right, here's a small ticket or here's something easy you can do.' And you don't have as much opportunity for really ambitious projects because all of the prototyping that I was doing was just really ambitious short projects that I would get very quick feedback on and then move on to the next one.
你知道,如果让我回想 2003 年,我当时在芝加哥一家披萨店工作,我的工作是烤披萨然后切好,对吧?然后放进盒子准备好。在芝加哥,你可以把披萨切成方块,对吧?我不知道。你那里也这样吗?
You know, if I cast myself back to 2003, I was working at a pizza place in Chicago and my job was to cook the pizzas and cut them, right? And put them in a box and put the box ready. In Chicago, you can cut a pizza square cut, right? I don't know. Do they do that where you're from?
是的。方块切法就是一切。
They do. Square cut was everything.
这里也一样,对吧?但有时候,你知道,那些没文化的人会点扇形切的披萨,所以我们就得那么切。很容易搞错,因为你习惯了某一种切法。但这是一个特殊要求。嘿,扇形切,你知道吗?
Same here, right? But sometimes, you know, these uncultured people would order pie cut pizzas and so we'd have to do that. It was so easy to get it wrong because you just get used to cutting it one way or another. But it was like a special request. Hey, pie cut, you know?
哦天哪。好吧。我们有一次接到一个要求,要一个扇形切的披萨,切成 11 片。你能切成 11 等份吗?我看着这个披萨,心想,我不知道。我不知道怎么切。
Oh god. Okay. Well, we got this request once for a pie cut pizza with 11 slices. Could you cut it into 11 equal slices? I'm looking at this pizza and I'm like, I don't know. I don't know how to do this.
是啊,拿出你的量角器。我叫了别人,他们过来,也说不知道怎么切。突然我们所有厨师,整个办公室都围着这些披萨,心想,等等,你会吗?不,不,不能那样。必须等分。然后,你知道,浪费了三张披萨后,我们碰巧切对了。我一直在想,是啊,你会怎么做?一定有办法切。多年后,我在 Quora 上问了这个问题,大概是 2008 年 Quora 刚出来的时候。我说,嘿,这是我遇到的事。你们都是 Quora 上的数学高手,你们会怎么做?这成了上面最热门的问题之一。
Yeah, get your protractor out. Yeah, I called someone else, you know, they come over, they're like, I don't know how to do this either. And suddenly we have all the cooks, the whole office was sitting around these pizzas being like, wait, would you? No, no, it can't be that. It has to be equal. And we just, you know, three pizzas later, we managed to get this thing just by luck. And it always was like, yeah, how would you have done that? There must have been a way to cut it. Years later, I asked this question on Quora when Quora was out like in 2008. I was like, hey, here's the thing that happened to me. How, you know, you guys are all math people on Quora. It was early Quora when it was good. How would you have done this? And it became one of the most popular questions on there.
是啊。
Yeah.
所以那个问题一直在我脑海里,关于这些小视觉技巧,视觉计算之类的东西。但直到很久以后,当我能够编程时,我才重新回到这些东西上。阴影怎么工作?样条曲线怎么工作?还有开源工作中的其他东西?这就是我开始研究的东西,因为都是视觉的。我在 Twitter 上发布,人们很喜欢,比如哦,这是怎么做等距的,或者怎么做样条曲线,等等。这主要发生在疫情期间。我当时在为 Play 做原型设计。那是一个很棒的团队,很棒的产品。
And so that question just sort of always lived in my head of these little visual tricks, kind of visual computation things. But it wasn't until much later when I could kind of code that I found my way back to this stuff. How do shadows work or how do splines work or these other things in the open source work? That's the type of thing that I started working on since it was all visual. I was posting on Twitter and people were liking it, like oh here's how to do isometric stuff or here's how to do splines or here's all this. This is mostly happening during COVID. I'm doing prototyping for Play. It was a great team, great product.
2020 年夏天,在 Framer 和 Play 之间,我还在做一个项目。我知道在 Framer 结束后,我还有几个月时间才会找到下一份工作。那时我迷上了状态机和状态图。所谓状态机,就是用户只能处于有限的状态,比如已登出或已登录。在每个状态下,用户触发某些事件,最终你会得到系统中所有状态和事件的列表。这是一种递归的方式来描述你在应用中的位置以及可能发生的事情。我一了解到这个,就觉得这对设计师来说太棒了,非常适合用来讨论一个东西是如何运作的。有点像乔布斯那句俗套的话:重要的不只是外观,还有它的运作方式。听起来很棒,但我没有工具来设计运作方式。于是我建了一个叫 State Designer 的应用。
I'd also worked on a project in the summer of 2020, during COVID, between Framer and Play. I knew I had a couple of months before I would find my next job after Framer ended. I'd become kind of obsessed with state machines and state charts. What I mean by state machines is there's only so many states that a user can be in, right? They can be logged out, they can be logged in. For each one of those states, they have certain events that can happen while they're in that state. And then at the end, you have a list of all the states and all the events in my system that can happen. So it's this recursive way of describing where you are in the application and what can happen. As soon as I learned about it, I thought, 'Oh, this would be so good for designers and so good for talking about how a thing works.' Kind of a cheesy Steve Jobs quote: it's not just what it looks like, it's also how it works. That sounds great, but I don't really have any tools for designing how it works. So I built an app called State Designer.
在开发那个应用时,我需要箭头,这些箭头要指向下一步。比如我点击登录事件,希望它指向图中的下一个节点。但问题是,这个图还需要是响应式的。回想起来,其实它不需要响应式,但我就是觉得应该这样。我不知道框会放在哪里,不能只画一次箭头就重复使用,因为位置可能会变。我无法提前知道元素的位置,所以需要想出一种在两个框之间画箭头的方法。你可以用很 naive 的方式,画一条直线箭头,但那会很难看。我想我必须做得更好。如果这里有一个框,那里有一个框,好看的箭头应该是什么样的?我会在纸上画:框在这里,框在那里,箭头会是什么样?然后我想,如果这个框更大,那个框更小呢?箭头会是什么样?如果两个箭头水平对齐呢?如果一个在另一个里面呢?于是我画了几百个框和箭头,尝试各种大小和位置的组合。然后我看着它们,觉得这里面有规律。让我试着把它形式化,用代码实现出来。结果这成了我做过的最复杂的代码:只是画这些完美的箭头。边缘情况多得离谱。天哪,但交点会丢失,因为一个在另一个里面,或者太近了。什么时候判定太近了?结果发现这非常主观。最后我只能说,这个问题没有答案,只是我觉得什么样的箭头好看。而我已经画了 20 年画,我知道好看的箭头是什么样的。当然,如果有人应该注意到这些,那不应该是我。
While I was working on that, I needed arrows, and I needed these arrows to point to what was going to be next. Like if I'm clicking on the sign-in event and I want it to go to the next thing in this graph. But the thing was that the graph also needed to be responsive. Thinking back to it, no, it didn't need to be responsive, but I just felt like it should be. I didn't know where the boxes were going to be. It wasn't like I could just draw an arrow once and reuse that arrow because things might be in a different place. I didn't know ahead of time where the things were going to be. So I needed to come up with a way to draw an arrow between these two boxes. You could do it in a really naive way, which would look terrible, just a straight arrow. But I thought, 'I've got to be able to do better than that.' If you have a box over here and a box over here, what would a good arrow look like? I would just draw it on paper: box here, box here, what would that arrow look like? Then I'd think, 'Okay, what if this box was bigger and this box was smaller? What would that arrow look like? What if the two arrows were horizontally aligned? What would that arrow look like? What if one is inside the other? What would that arrow look like?' So I'm just drawing hundreds of boxes and arrows and combinations of different sizes and placements. Then I looked at them and thought, 'I sense there's a pattern here. Let me try to formalize it. Let me try to code this up.' And it became the most complicated thing I'd ever tried to do with code: just trying to draw these perfect-looking arrows. The amount of edge cases was insane. 'Oh my god, but the intersection would be lost because the thing's inside of it, or it's too close. When do I decide that something's too close?' It turned out to be highly subjective. At the end, I was just like, 'Well, there is no answer to this. It's just what do I think is a good-looking arrow?' And I thought, 'I know what I've been painting and drawing for 20 years. I know what a good-looking arrow looks like. Of course, if anyone should notice this stuff, it shouldn't be me.'
我在 Twitter 上发帖,虽然我的受众很小,但有些帖子引起了很大兴趣。人们跟着看,说,瞧这家伙,他为箭头抓狂了。看看这个。我不得不做调试工具,它看起来像一个弹簧系统,箭头会弯曲绕过不同的边,选择路径,然后收缩和展开。最后我做出了一个效果不错的东西。它用在了我的 State Designer 应用里,但我也把它开源了,作为一个叫 Perfect Arrows 的库分享出去。这个库也很受欢迎,因为这个问题看似应该有答案,但对于开源项目来说,它很不寻常,因为没有标准答案。这是一个设计上的答案,通过代码表达出来。这是我对在两个物体之间画箭头这个问题的理解:这是所有参数,这是我在所有情况下让它看起来好看的最佳尝试。人们真的很喜欢。对于一个开发者来说,解决这类问题似乎很不寻常,也很新颖。很多开源项目是关于众所周知的问题:排序一个很长的列表,在数据类型之间转换。你知道参数,只需要实现它,让它可复用、轻量、快速。在状态好的时候,如果你自己动手,你会得到同样的代码。但对我来说,这就像是,你可能不会为箭头痴迷,但如果你痴迷了,而且你对箭头有很好的品味,那么这大概就是你会得到的结果。这已经远远偏离了开源通常的做法。我喜欢这样。我觉得,这太棒了。
I was posting about this on Twitter, and even though I had a fairly small audience, some of those posts were getting a lot of interest. People were following along like, 'Yo, check out this guy, he's losing his mind over arrows. Get a load of this.' I had to make debugging tools, and it looked like a spring system, collapsing and expanding as the arrows would bend around the different sides and pick which route to take. I ended up with something that worked pretty well. It went into my State Designer app, but I also open-sourced it and shared it as a library called Perfect Arrows. That got pretty popular as well because it was a problem that seemed like it should have an answer, but for an open source project, it was very unusual because there was no answer. It was a design answer, expressed through code. This was my educated take on the problem of drawing an arrow between two things: here are all the parameters, here is my best shot at making it look good in all circumstances. People really liked that. It seemed unusual and new for a developer to be tackling that type of problem. A lot of open source is about well-known problems: sorting a very long list, converting between data types. You know the parameters, it's just about implementing it in a reusable, lightweight, fast way. On a good day, if you did it yourself, you'd end up with the same code. But for me, it was like, 'Here's this thing where you probably wouldn't obsess over arrows, but if you did and you had good taste in arrows, this is probably what you'd end up with.' That's already so far from the normal way open source is done. I liked that. I thought, 'Oh, this is great.'
另外,还记得我在 Framer 的那份工作吗?当时我试图为 Framer 制作人们真正喜欢的内容。我的职位是设计教育者。我试图制作关于 Framer 产品的内容,但必须是吸引设计师的技术内容。那时 Framer 更偏技术,代码很多,我搞不明白。教程视频?没人看。博客文章?没人读。实际上没人真正想学。我觉得这是部分原因。是有原因的。但我仍然在想,为什么我不能做出好的内容呢?我知道那个受众是存在的。
Also, hey, remember that job I had at Framer where I was trying to make content that people really liked for Framer? That was my job at the time: design educator. I was trying to make content about the Framer product, but technical content that appealed to designers. At the time, Framer was much more technical and code-heavy, and I couldn't figure it out. Tutorial videos? No one would watch them. Blog posts? No one would read them. No one actually really wanted to learn. I think that was part of it. There's a reason. But I was still wondering, 'Why couldn't I just make good content? I know that audience exists.'
我知道,当我去伦敦参加 Framer 聚会时,现场大约有 200 人,我问“有多少人以前用过 Framer?”大概只有四个人举手。大家对这类内容有兴趣,而且是一种非常向往的兴趣。你怎么才能做出那些有这种兴趣的人会在乎的内容呢?我觉得这些箭头之类的东西越来越接近了。感觉人们不需要读任何代码,他们就能看到问题的复杂性,看到我做的决策,而我可以用一条简短的 Twitter 推文配上 GIF 轻松地讲清楚。那是很棒的内容。所以我想继续做下去。
I know that when I went to Framer meetups in London, there'd be like 200 people there and I'd say, 'How many people have used Framer before?' and like four people would raise their hands. There was just an interest in this type of content and a very aspirational interest as well. How do you make content that people with those types of interests would care about? And it felt like this arrow stuff was getting closer. It felt like people didn't have to read any code. They could kind of see the complexity of the problem and see the decisions that I was making, and I could talk about that really easily in a short Twitter thread with GIFs. It was great content. So I wanted to keep going.
接下来的一个小瞬间、小片段或故事章节是关于 Play 这个移动应用的。我想录制一些关于 Play 的视频,比如如何使用它。我可以把手机连到电脑上,然后接入 OBS,这样大家就能看到我的屏幕。但你看不到的是,我在制作 Framer 教程视频时非常依赖的一点:在桌面上你有一个光标。当你共享屏幕时,光标就像是你的手,你用这个光标引导观众的视线。但在这种情况下,我做不到,因为我的电脑连着手机,我没有带手指的摄像头。我不能触摸屏幕,因为如果触摸屏幕,那就会对系统产生实际输入。我想,我真正想做的是在屏幕上方画图,就像美式足球比赛里画那些又大又粗的黄色箭头一样。于是我做了,而且效果非常酷。
The next kind of little moment or vignette or chapter of that story was Play, a mobile app. I wanted to record some videos about Play, like how to use it. I could connect my phone to my computer and put that into OBS so that you could see my screen. But what you couldn't see, and what I had really relied on when making tutorial videos for Framer, was that on the desktop you have a cursor. When you're sharing your screen, the cursor kind of is your hand. You guide the eye of the viewer using this cursor. In this world, I couldn't do that because my computer was connected to my phone and I didn't have a camera with my finger. I couldn't touch the screen because if I touched the screen, that would be actual inputs on the system. I thought, you know what I really want to be able to do is just draw on top of my screen, kind of like an American football game where you're doing these big fat yellow arrows. So I did it, and it was really cool.
顺便说一句,关于在屏幕上画图,我学到了很多。这叫做 Telustrator,是 60 年代发明的,用来在屏幕上画图。我想,“好吧,我要做一个 Telustrator。”然后我做了。它非常棒。这是一个 Electron 应用,位于屏幕前方但透明,允许事件穿透。但如果你拨动开关或按下键盘快捷键,它就会开始捕获事件,并用它来运行这个绘图工具。画出来的东西会逐渐淡出。这正是我想要的。它非常酷,我用上了。太棒了。
By the way, I learned a lot about drawing on top of your screen. It's called the Telustrator, invented in the '60s to draw on top of your screen. I'm like, 'All right, I'm going to make a Telustrator.' And I did. It was really nice. It was an Electron app that was in front of your screen but transparent and would allow events to pass through. But if you flipped a switch or did a keyboard shortcut, it would start capturing events and use that to run this drawing tool. The drawings would kind of fade away. It was exactly what I wanted. It was really cool and I used it. It was great.
然后我想,这很酷,但我用的手写笔是 Wacom 的数位绘画笔,是我搞艺术时买的,挺贵、很好用的笔。它能感知压力,我在 Photoshop 里用过,所以我知道。我就想,怎么把它用进我的小 Illustrator 应用里呢?我发现如何从指针事件中获取压力值,这很容易。但如何根据压力画出一条粗细变化的线呢?没有答案。我深入挖掘了一下。有一个 React Native 的签名组件。它是怎么工作的?效果很糟糕。这是一个很有野心的项目,但我绝对不会发布它。这显然不是答案。肯定有更好的办法。它通过绘制大量小线段并改变每条线的宽度来实现。在浏览器中,你没有 SVG 图元能支持一条线沿长度方向改变宽度。
Then I thought, this is cool, but the stylus I have, because it was a digital painting stylus from Wacom, it was kind of an expensive, really nice stylus from my art days. It could do pressure. I knew that because I could use it in Photoshop. So I'm like, how do I get that into my little Illustrator app? I found out how to get the pressure off a pointer event. That's really easy. But how do you make a line that gets bigger and smaller based on pressure? No answer. I dug a little deeper. There was a signature thing for React Native. How does it work? It works terribly. It was a good ambitious project, but I would never have shipped this. It's clearly not the answer. There's got to be a better thing. It does it by making lots of little tiny lines and varying the width of each line. In the browser, you don't have SVG primitives for a line whose width changes along its length.
所以你必须伪造它。
So you have to fake it.
你可以用线段来伪造,有宽段、窄段、更窄段。或者你可以用 Photoshop 的做法,叫做“点画笔刷”,取一个形状,比如圆形,然后重复很多很多次,根据压力变大或变小。因为形状靠得很近,最终看起来像一个连续的形状。这两种方法的问题在于它们只适用于光栅图像,只适用于像素。我想,我要在 SVG 里做这个。我不想要那些看起来也不好看的小形状和线条。我想要一个多边形,比如用点来包裹我画的所有点,构成一个矢量形状。结果发现没人真正搞明白这个。我决定自己来解决。
You can fake it with line segments that have wider segment, narrower segment, narrower segment. Or you can do the thing that Photoshop does, called a dab brush, where you take a single shape, say a circle, and repeat that same shape lots of times, getting smaller or bigger depending on the pressure. Because the shapes are so close together, you end up with what looks like a single continuous shape. The only problem with both of those things was that they only worked with raster images, only in pixels. I was like, well, I want to do this in an SVG. I don't want little shapes and lines that don't look good anyway. I want a polygon, like points that wrap over whatever the points that I drew and constitute a vector shape. Turns out no one had really figured that out. I took it upon myself to figure it out.
我开始做这个是因为我刚看了一个关于如何制作赛道的视频。如果你在编程一个复古电子游戏,想要一条赛道,你会取所有中间点,然后向左和向右扩展,这样就做出了赛道。但我想,如果你在压力更大时向两侧扩展得更远,压力更小时向两侧扩展得更少,那就能画出一条粗细变化的线。我可以搞明白。但当我真正开始做时,发现这其实是个非常难的问题。
I started this because I just watched a video about how you do racetracks. If you were programming a retro video game and wanted a racetrack, you take all the middle points and go out left and out right, and that's how you make the racetrack. But I thought, if you just went further to the sides, further left and further right when there was more pressure, and less to the left and less to the right when there was less pressure, that would create a line that got thicker and thinner. I could figure that out. But I jumped into it and that turned out to be a really hard problem.
我就知道你要说这个。
I knew that was about to be the next thing you said.
我在 Twitter 上发帖,分享这些小演示和多边形骨架的调试视图。有人私信或回复我说,“哦,我博士期间就研究这个,我的整个博士论文都是关于拐角的。”我想,哦,我完全搞不定了。我和一些人通了电话,他们说,“不,你需要学线性代数,伙计。如果你想做好这个,你得搞明白这些东西。”于是我学了。向量、二维向量、均匀大小。我学会了,因为我想做的是墨水。所以不管学什么才能做出墨水,我都会去学。
I'm posting about this on Twitter. I'm sharing these little demos and debug mode views of skeletons of these polygons. I'm getting people in my DMs or replies being like, 'Oh yeah, I worked on this for my PhD and my whole PhD was just about the corners.' I'm like, oh, I'm way over my head. I had calls and talked to folks who were like, 'No, you need to learn linear algebra, buddy. If you want to do this right, you're going to need to figure this stuff out.' And so I did. Vectors and 2D vectors and uniform magnitude. I got it because what I wanted to do was make the ink. So whatever I got to learn to make the ink, I'll figure that out.
然后你又在 Twitter 上公开地、近乎痴迷地研究这个问题好几个月,以至于大家都说“哥们,你就认了吧”,而你会说“不不不,还有好多东西可以尝试和学习,我连皮毛都没摸到呢”。我坐在这儿脸上带着笑,因为我觉得我刚刚更了解了你一点:首先,你对那些大多数人绝不会去碰的、极其困难的问题有着强烈的渴望。
And then just again like kind of obsessively working this problem for months on Twitter in public to the point where like folks were just like dude you gotta just call it you know like you got you know and be like no no no there's still so much other stuff to try and to learn or to you know I haven't even scratched the surface or whatever. I'm like sitting here with a smile on my face because I feel like I just got to know you a little bit where one you have this appetite for absurdly hard problems that most people would never go down that rabbit hole.
你显然学得很快,但又对细节有执念,愿意从第一性原理出发拆解每一个小部分,以找到构建东西的最佳方式。再加上你对教学、教育和解释想法的兴趣,而且你意识到这在很多方面可能是最好的学习方式。所以,综合所有这些,你在工具设计上深耕了很多年。那么,如果你要开设一个关于工具设计的系列内容,你会强调哪些主要原则,来帮助听众中的设计师真正深入那个世界,理解你必须思考的那些事情?
You also very clearly are fast learner but then you have this obsession with the details and a willingness to tease apart every little piece from first principles in order to figure out the best way to build something. So take in that experience and then also there's some interest in teaching and educating and explaining ideas and you're able to recognize the fact that that's also in many ways probably the best way to learn something. So if we combine all of that you've spent years in tool design like really really deep in tool design. So, if you were teaching some kind of an upcoming content series on tool design, what are some of the main principles that you would be hitting in order to help a designer listening really go deep into that world and understand just the types of things that you have to think through.
哦,是的。我认为所有工具本质上都是决策工具。比如颜色选择器就是一个很好的工具,对吧?它不是表示颜色的唯一方式。你可以手动用十六进制代码或 RGB 代码来选颜色,但那很糟糕。那不是做决策的好方式。你无法微调,无法快速获得关于某个决策的反馈,也无法进行比较。一个好的工具能让你安全地做出更改并回退,给你一个安全网,让你在操作时不会破坏任何东西或搞砸事情,这也是我觉得设计系统有点有趣的原因。总之,你需要那种安全感来做决策。你还需要能够比较选项,能够创建这些选项。这需要大量的微调,需要在精确性(能够精确操作)和缺乏精确性(以便做出广泛的决策)之间取得平衡。换句话说,如果应用只允许你调整间距,并且对间距有很好的控制,那很好,但那不是我在这里需要做的全部。如果你不得不把每个决策都做成精确决策,那么操作速度就会变得非常慢,甚至让你脑子装不下,因为你必须做太多决策了。甚至一些不重要的事情你也得做决策。而一个好的工具能让你专注于你正在做的决策或工作,无论它是否精确,而不会让你负担所有其他创造性决策。
Oh, yeah. I think of all tools as like kind of decision-making tools. Like a color picker is a great tool, right? Not the only way of representing color. like you could pick a color like using like hex codes or RGB codes or something manually. It's just that's bad. It's not a good way of making those decisions. You're not dialing anything in. You're not getting that like really quick feedback on one decision or another and being able to compare things. A good tool will allow you to to do that to to very safely make a change and go back have that safety net of I'm not destroying anything or I'm not like screwing anything up while I'm working on this, which is why I think design systems kind of funny. Anyway, yeah, like you know, you need that like that that that safety in order to to to make those decisions. You also need to just be able to compare options. You need to be able to make those options. It's a lot of dialing in. That's a lot of uh balance between precision in terms of the being able to to do things precisely and then a lack of precision in in other ways in order to allow you to make a breadth of decisions. In other words, like if if the app only allows you to just do the gaps, you know, but it gives you a really good control over the gaps, like that's great, but like that's not the only thing that I need to do here. If you have no choice but to make every decision a precision decision, then like the speed of of operating just becomes really slow or even like kind of too much to to hold in your head at the same time because like you're having to make too many decisions, I suppose. uh even things that aren't important you might have to make decisions on. Whereas a good tool will allow you to kind of focus on the decision that you're making or the work that you're doing which may or may not be precise without burdening you with having to make all those other creative decisions.
你显然对好工具很有眼光,能看到所有细节和成千上万的微决策。所以,当你使用这些不同的基于画布的工具时,我敢肯定你经常会有“嗯,这个可以改进”或者“我能看到一种不同的做法”的想法。那么我想深入探讨这种张力:你如何平衡创新的欲望和利用现有熟悉度的需求?即使你能看到更好的方法,但完全匹配现有做法也有真正的价值。作为工具设计师,作为像 tldraw 这样在基础层工作的人,你如何看待这种张力?
You obviously have an eye for what good tools are and you can see all of the details and the thousands of micro decisions. So as you're using these different canvas based tools, I'm sure you get to the point semi-frequently where you're just like well that could be improved. Well, I could, you know, I could see a way to do that a little bit different. So then I want to dig into this tension where it's like, how do you wrestle with one the desire to innovate and improve while simultaneously capitalizing on the familiarity that exists and you know the real value in matching something one-to-one even if you can see a better way to do it. So how do you think about that tension as a tool designer and someone that's kind of working on that foundational layer like tldraw?
这东西之所以能成为商品或开源产品,是因为画布是一种已知的东西。它是一种商品,或者至少理论上是这样:当你使用画布产品时,你会自动带入大量从 Figma、Miro 或其他工具中习得的可供性,并期望它按预期工作。你不会对每个软件都抱有同样的期望,认为工具是已知的东西,对吧?但画布的“已知”非常复杂。
The reason why this can be a commodity product or is an open source product is because the canvas is kind of a known thing. It's it is a commodity or at least that's the theory is that like when you use a canvas product you you automatically kind of bring with you tons and tons and tons of affordances that you've picked up by using Figma or by using Miro or or any of these other tools and you expect it just to work the way that it's supposed to. You don't bring those same type of expectations to every piece of software of like the tool has a kind of like it's a known thing, right? But the the known thing for canvas is very complicated.
是的。
Yeah.
我举另一个高度商品化但又极其复杂的东西的例子:文本编辑器。如果你在文本编辑器里打字,停了一下,然后继续打,接着你说“啊,不对”,然后按撤销。如果你按撤销,撤销的是你刚打的最后一个字符,再按撤销又只回退一个字符,你会觉得这东西坏了,对吧?因为它应该回退到你暂停的地方,对吧?反正这就是它的工作方式,这就是惯例。文本编辑器中撤销/重做有非常强的惯例。如果不按那样工作,你不会说“哦,这个文本编辑器只是有点不同,我得按很多次才能回到我想去的地方”,而是会说“我无法使用这个”。这简直是 IRS 级别的网络恐怖。我是说,它完全不可用。即使其他功能都正常,只要一个重要的惯例化特性搞错了,你就会觉得“这根本不是文本编辑器,对吧?”画布也有同样的事情。比如你在画布上双指捏合,应该放大或缩小,但缩放中心不应该是屏幕中心,而应该是你捏合的位置。如果缩放时不是这样,就会感觉坏了,感觉这只是一个技术演示,不是真正的产品。如果你选中一堆框然后旋转,它们应该一起旋转。这让我作为工具制造者必须认识到,在画布这样的东西里,成千上万个特性中哪些需要每次都保持一致。
The example I use for another one of these like really well commodified and yet super complex things is like a text editor, right? So, if you're typing into a text editor and you kind of like pause and they keep typing and then you say, "Ah, that's not right." And you you hit undo. If you hit undo and the thing that was undone was the last character that you typed just and you did undo again and went like just one character back, you'd be like this thing is just broken, right? Because it should go back to where you paused, right? It could be anyway, but like that's the way that it works, right? That's the the convention. very strong convention around undo redo in text editors. If it doesn't work that way, it's not like, oh well, this text editor just works a little different. You know, I got to press this a lot of times in order to get back to where I want to go. It's like, I cannot use this. That's like, you know, uh IRS level um you know, web terror. I mean, it's uh it's just completely unusable. Even if the rest of it all works, like if you get one of those really important conventionalized features wrong, like it just it's like, well, this not a text editor, is it? You know, the canvas has the same sort of things that should happen. Like if you pinch on the canvas, the thing should zoom in or should zoom out, but it shouldn't zoom out towards the center of the screen. It should zoom in to where you're pinching. And if you're zooming out, it should zoom out from that. That should be the origin of the transformation of the canvas. And if it doesn't work like that, it feels broken. It feels like, oh, these is like a tech demo. This is not a a real thing, right? If you select a bunch of boxes and you you rotate them, they should all rotate together. That does put me as a like a toolmaker in a in a spot of having to recognize what of the thousands and thousands and thousands of features inside of something like a canvas needs to be the same way every time.
然后还要识别,嗯,哪些实际上是不一样的,对吧?比如对齐或两端对齐在不同应用里是怎么工作的,有没有原因?我注意到,我们(或者我)在 geraw 上的很多决定会越来越影响这些规范,因为总有一天会有人说,好吧,ClickUp 在这种情况下怎么做?Inflight 呢?Autodust 呢?然后发现,哦,它们全都一样。
And then also identifying, okay, which ones are actually different here, right? Like how does alignment or justification work that is different between these different apps and is there a reason for that? I do notice the fact that like a lot of our decisions with geraw or a lot of like my decisions with geraw are going to be increasingly influential on what those norms are because some point someone's going to say like okay well what does clickup do in this case all right cool uh what does inflight do in this case you know and like what what is uh you know these what does autodust do in these cases and like like oh look they all do it the same way.
嘿,很快说一下全新的 Dive 人才网络。我亲手组建了一个超过百人的设计师和开发者人才库,都是我知道的最有才华的人,这样我就可以把他们推荐给我最喜欢的公司。所以,如果你在听这个节目,并且对新的机会持开放态度,这个人才网络是匿名的,而且超级轻松。这只是看看外面有什么机会的简单方式,不用在社交媒体上发帖。所以,如果你有兴趣加入,或者正在寻找下一个员工,请访问 dive.comclub/talent。
Hey really quickly let me tell you about the all new dive talent network. I've hand assembled over a hundred of the most talented designers and builders that I know so I can recommend them to my favorite companies. So, if you're listening to this and you're open to new opportunities, the talent network is anonymous and super low pressure. It's just an easy way to see what's out there without having to post on social media. So, if you're interested in joining or maybe you're looking for your next hire, head to dive.comclub/talent.
这就像你举起了一个绝大多数人都没有的放大镜。你看到的细节层次是独一无二的。
It's almost like you hold up a magnifying glass that the vast majority of people do not have access to. Like you're seeing things at a level of detail that is unique.
我想稍微拉远一点,因为我在做更多演示时看到了一个有趣的趋势:画布作为下一代工具范式的趋势已经很明显了。确实有增长。显然,现在是做画布 SDK 业务的好时机。但你上次我们聊的时候说了一些很有意思的话。你谈到今天大多数应用对技术的使用仍然很保守。你能稍微展开讲讲吗?我最终想帮助人们稍微进入你的大脑,看看你想象中这能走向何方,并拓展人们设想未来的能力。
I want to zoom out for a second because there's an interesting trend that even I'm seeing as I do more of these demos where the canvas as a paradigm for this next era of tooling is quite clear. Like there is an uptake. It's a good time to be in the canvas SDK business apparently. But you said something last time we talked that thought that was pretty interesting. You talked about how most of today's apps are still conservative uses of the technology. Can you unpack that a little bit? And I want to ultimately help people get inside your brain a little bit where you're imagining where this could go and stretching people's ability to envision the future.
我认为仅仅在设计工具这个类别里,设计能做的事情就多得多。我觉得你也在做其中一些,比如这是一个用于创建生产素材的环境,或者这是一个用于创建网站或应用的环境,或者这是一个设计可以以某种方式执行的环境,有点像 Taildrop 计算机或者我们看到的一些工作流工具,或者这是一个映射回真实物理过程的设计。比如我见过在制药领域使用画布的工具,那里有非常复杂的研究或生产工作流需要设计,但你只能设计尚未运行的部分。所以这个工作流就像时间在向前移动,后面的一切你都无法再改变,因为它已经发生了,但前面的一切你都可以改变,并且可以从这里学习以便改变。很迷人,对吧?考虑到复杂性和所有这些不同事物之间的关系,最简单的方法就是可视化地去做,在一个可以缩放和平移的地方去做。
I think within just the category of design tool, there's a lot more that you can do in design. I think you're doing some of this as well like saying this is an environment for creating production assets or this is an environment for creating websites or applications or this is an environment where the design is executable in some way kind of like Taildrop computer or some of the workflow tools that we're seeing or this is a design that is mapping back to a real physical process even. Like I've seen tools that use the canvas in pharmaceutical for example where you have these really complex workflows either doing research or production that need to be designed but you can only design the parts that haven't run yet. So it's like this workflow and it's like time is moving along this and everything back here you can't do anything about anymore because it happened but everything this way you can change you know and like learn from here in order to change. Fascinating right? The way to do that easiest given the complexity and given the relations of all these different things is to do that visually and to do that in a way place where you can zoom in and out and move things around.
我认为工作流本身现在仍然很流行,但也是一个不断扩展的类别,你可以用工作流做更多事情。我把工作流看作是画布的一个子集,属于白板类别,这个类别已经很知名了。但当你开始深入这些垂直领域时,比如我认为 Miro,人们用它做 UX 研究、团队会议、教育、远程异地会议、入职培训等等,Miro 能容纳所有这些用途,我认为这是它的优势,但同时也是一种不太适合所有用途的情况。你知道,就像你说,如果 Miro 只做团队入职活动或远程异地会议,它会是什么样子?功能集可能会少很多,首先,但也可能会有一些独特的东西。你在画布上表现人的方式可能会不同,或者更……我的意思是,我不知道,因为我不是在构建那个产品,但我看到的是,人们对更深入地开发这些不同垂直领域有兴趣,最终它们看起来不太像白板。它们看起来像 Padlet 之类的东西,基本上是在重建 Hyperard,因为他们想给老师提供这些演示工具,让学生们去探索。我见过课堂工具,甚至有一些 AI 功能,画布上会根据学生的问题或他们的研究自动生成内容。这仍然是一种可视化呈现信息的方式。同样的画布。
I think workflows themselves are still very popular right now but also kind of expanding category of what you can do with workflows, and I consider workflows a kind of a subset of canvas within like kind of the whiteboarding category that's pretty well known. But you start kind of going into these verticals right where I think Miro for example, like people use that for UX research, for team meetings, they use it for education, for remote offsites, for onboarding people, for so many different use cases in Miro, which I think is a strength for Miro that it can accommodate all those things, but it's also kind of a little bit of like a bad fit for all of them. You know, like you say like, oh, what would Miro look like if the only thing that it did was facilitate team onboarding events or remote offsites or something like that? What would the feature set be? Would probably be a lot less, number one, but also it would probably have things in there that are unique. The way that you represented people on the canvas, you know, might be different or might be more... I mean, I don't know because I'm not building that product, but what I've seen is that there's interest in developing those different verticals much deeper in ways that eventually don't really look like whiteboards. They look like things like Padlet, you know, where it's like they're rebuilding Hyperard essentially because they wanted to give these presentation tools to teachers to build these environments for their students to explore. I've seen classroom tools where there's even some of the AI stuff coming in where you have auto content that is being generated on the canvas in response to the questions from the students, in response to the research that they're doing. It's just a visual way of presenting that information. Again, same canvas.
我认为多人画布体验非常好,而且被严重低估了。我甚至觉得,比如……我一直在构建很多入门套件,下一个我想创建的入门套件就是一个棋盘游戏。比如我想在画布上放一个骰子,放一副牌。我想看看人们能用它构建什么,因为这是一个多人实时画布。现在我手里有骰子了。好吧,我们来玩龙与地下城吧。我们来玩扑克吧。你拥有所有基本元素:我需要选择、移动、拖拽、激活。我有协作者,他们有不同的位置和拖放区域,等等。所以很多完整的产品类别都可以融入画布,但我们熟悉的那些仍然是非常早期的版本。
I think the multiplayer canvas experience is so good and so underexplored. And I think even things like... I've been building a lot of starter kits and one of the starter kits that I want to create next is just a board game. Like I want to have a dice on the canvas. I want to have a card deck of cards on the canvas. And I want to see what people build with that because it's a multiplayer real-time canvas. And now I have dice in my hands. Like, all right, let's play some D&D, man. Like, let's play some poker. Like, you have all the primitives there. I need to select, I need to move, I need to drag, I need to activate. I have collaborators and they have their different spots and drag and drop areas, all that stuff. So it's like so many entire product categories can fit into the canvas, but the ones that we're familiar with are again like it's a really early generation.
就在我们录制这期节目的一周前,Figma 宣布收购 Weevi。光是这件事就会让很多设计师对画布能做什么以及这些改变工作流的方式产生不同的思考。我们在开始录制之前也聊过这个。人们带着演示来找我,我看到了很多基于画布的演示。
Even just seeing like we're recording this a week after Figma announced the Weevi acquisition. Like even just that is going to introduce so many designers to a different way of thinking about what a canvas can do and these change workflows. And we were talking about this before we hit record. Like people come to me with demos and I'm seeing a lot of canvas based demos.
我越来越清楚地感觉到,我未来很多编码工作可能会扎根于画布——在画布上并排查看不同版本的本地主机,比较、对比、创建分支,一切都在画布上。感觉未来会更像这样,挺令人兴奋的。
It's starting to feel clear to me that a lot of the coding that I'm going to do might be rooted in the canvas where I'm seeing different versions of my local host side by side and comparing and contrasting and creating branches and everything exists on a canvas. It's like whoa, more of this future feels like it's going to exist on the canvas in a way that is pretty exciting.
回到什么造就了好工具这个话题,本质上代码在构思阶段并不擅长。它不擅长做决策和比较事物。虽然有办法做到,甚至现在一些 AI 工具如 Cursor 有多个智能体在多个工作树中工作,你可以切换——但还是很笨拙。不是产品不够天才,而是这根本不是比较事物的方式,也不是微调和调试的方式。
To go back to what makes a good tool, essentially code is really not good at that ideation stage. It's really not good at making decisions and comparing things together. There are ways to do it, and even some of the AI tools out right now like Cursor with multiple agents working at multiple work trees so you can switch between them—it's all very clunky still. Not for lack of genius of the products, but it's just not the way you compare things. It's not the way you tweak and dial in and all those things.
在画布上你可以做到。人们为什么在设计领域如此喜欢画布工具,这不是偶然的。拥有无限空间来比较、分支、不断深入、回退——所有这些在画布上比在 git 历史中容易得多。不是不可能,但难如登天。如果你要做,你做的次数会远少于在更适合迭代的环境中能做的次数。这太重要了。
You can do that on the canvas. It's not an accident why people really like canvas tools within the design space. This idea of having an infinite space to compare, branch, go deeper and deeper, rewind—all that stuff is way easier to do on the canvas than in a git history or something like that. It's not impossible, but prohibitively hard. If you're going to do it, you'll do it a fraction of the amount of times that you could do it in an environment more geared towards that number of iterations. That's such a big thing.
你知道,在我做决定之前,我能快速做出多少个不同版本?
And you know, how many different versions of this thing can I whip out before I decide to make a decision?
我们能聊聊你的实验吗?你总在 Twitter 上发这些小东西。我最近看到一些仙女在飘。
Can we talk a little bit about your experiments? You're always posting these little things on Twitter. I've seen some fairies floating around recently.
嗯,我现在在里斯本的万豪酒店,刚刚在里斯本 AI 大会上展示了仙女。跟你说,效果非常好。
Well, I am right now I'm in the Marriott in Lisbon where I just presented fairies at the Lisbon AI conference. And let me tell you, it went really well.
不错。
Nice.
我跟你说说。我们卖 SDK,授权给其他公司。我们的价值主张是你可以用 Tldraw 构建很酷且非常不同类型的东西。所以我有责任去构建很酷且非常不同类型的东西,对吧?而且我在 Tldraw 之前的那段日子里也学到,营销这类工具最好的方式就是用它来构建。用它构建有趣的东西,并在公开场合逐步开发。那些未解决的问题——明显没完成甚至可能无法完成的事情——对构建者受众来说最有趣,对吧?他们想跟着看别人在做什么,也想知道自己会怎么做。我觉得对我来说是这样,对大多数开发者和设计师也是如此。
I'll tell you about it. We sell the SDK, we license it to other companies. Our value proposition is that you can use Tldraw to build cool things and very different types of things. So it's kind of incumbent on me to build cool things and very different types of things, right? And I also learned during those days before Tldraw that the best way to market a tool like this is you just build with it. You build interesting things with it and you develop it a little bit in public. And the unanswered questions—things that are obviously not done or maybe even not doable—are most interesting to a builder audience, right? They want to follow along as someone's doing something and also wonder about how they would do it. I think that's true for me and for most developers and designers.
所以,我们用 Tldraw 构建了很多东西,但真正病毒式传播、几乎凭空出现的,是 Tldraw 加 AI 的各种方式。我快速过一下,因为这完全是另一期播客了,对吧?但我们确实做了 Make Real,你可以画一个网站或画点什么,选中它,点击一个叫 Make Real 的按钮,它就会创建一个网站并放到画布上,因为我们的画布可以做到。它会是你画的任何东西。我们截取那个截图,发送给 AI,然后说:“嘿,你是一个 4000 岁的高级网页开发者,热爱你的设计师,希望他们开心。你的设计师刚给了你这个低保真线框图,要求一个可工作的原型。你能构建出来吗?”然后我们就做了。那是 2023 年,视觉模型刚出来。实际上是一个叫 Sawyer Hood 的 Figma 设计师——顺便说 Sawyer 很棒——他基本用 Tldraw 做了原型,然后它开始病毒式传播。然后我们接手并推进了它。我们说:“好,让我们把它放回画布上,而不是放在某个模态窗口里。”
So yeah, the things we've built with Tldraw are a lot, but the things that really went viral and came out of nowhere basically were Tldraw plus AI in various ways. I'll speedrun this because it's a whole different podcast, right? But we did make Make Real, where you could draw a website or draw something, select it, and click a button called Make Real, and it would create a website and put it on the canvas because our canvas can do that. It would be whatever you drew. We would just take that screenshot, send it to an AI, and say, 'Hey, you're a 4,000-year-old senior web developer who loves their designers and wants them to be happy. Your designers just gave you this low-fidelity wireframe and asked for a working prototype. Can you build that?' And we did it. It was 2023, the vision models had just come out. It was actually a designer named Sawyer Hood at Figma—Sawyer's awesome by the way—he had basically prototyped that using Tldraw, and it started going viral. Then we took it and ran with it. We said, 'All right, let's put this back on the canvas instead of in some modal window.'
我还记得那条推文,第一次刷到的时候,我心想:“天哪。”
I remember that tweet still, scrolling and seeing that for the first time, and I was like, 'Oh my gosh.'
它霸占了 Twitter 好几天。
It took over Twitter for a couple of days.
然后我们不断发现新东西。我们发现如果你在网站顶部画东西,然后我们把那个作为下一个提示发送,哪怕只是一个空白白框,并说:“嘿,这个白框包含你之前做的网站。这是你上次做的代码。根据用户选择的内容,生成一个新提示。”如果你只是划掉了一个图标,它不知怎么就能理解,会说:“哦,对,那在右上角,一定是划掉了菜单图标。好,在下一次迭代中,我就去掉菜单图标。”然后你就得到了下一个可工作的原型,又是通过画图创建的。然后人们意识到你可以把截图放在旁边,说:“嘿,让它看起来像 stripe.com,”或者“这是一张某个图标的图片,做出这个图标。”然后一个接一个地提示,你可以给它像 UX 文档中的图示,比如“这个旋转拨盘就应该这样工作”,它就会说:“好,我能做到。”这太迷人了。我们实时学习着这一切。这发生在 vibe coding 应用之类的东西之前。对很多人来说,这是他们第一次做出能用的东西——产生任何形式的软件制品。简直疯狂。这些推文有数百万、数十亿的浏览量和互动。我的电话响个不停,我们读 Twitter 读到手腕都疼了。整整九天,基于这个“画个网站”的想法病毒式传播。
And then we kept on discovering new things. We discovered that if you drew on top of the website and we sent that as the next prompt, even just with an empty white box, and said, 'Hey, this white box contains your previous website that you made. Here's the code you made last time. Make a new prompt, using whatever the user has selected.' And if you had just kind of crossed out an icon, somehow it would connect the dots, be like, 'Oh, yeah, that's in the top right corner. It must be crossing out the menu icon. All right, in my next iteration, I'll just get rid of the menu icon.' And then you had the next working prototype, which you had created again just by drawing. Then people realized you could just put screenshots next to it and say, 'Hey, make this look like stripe.com,' or 'Here's a picture of a certain icon, make this icon.' And just next prompt after next prompt, you could give it figures like you'd see in UX documentation, like 'This is exactly how this rotational dial should work,' and it would be like, 'All right, I can do that.' It was fascinating. We were learning this in real time. This was pre-vibe coding apps and stuff. For a lot of people, this was their first time making something that worked—producing an artifact of software in any shape. It was madness. There were millions and billions of views on these tweets and engagements. My phone was ringing off the hook, and we were all getting RSI from reading Twitter. It was nine solid days of virality based on this draw-a-website idea.
在那之后,有些事情变得容易多了。融资变得容易多了。
After that, some things got a lot easier. Fundraising got a lot easier.
但这也是这个画布的完美用例。如果我没有已经在构建 Tldraw,一旦视觉模型发布,我就会开始构建 Tldraw。
But also it was like this is the perfect use case for this canvas. If I hadn't already been building Tldraw, once the vision model dropped I would have started building Tldraw.
为了利用这项奇妙的技术,你需要一个真正好用的、可 hack 的画布来处理图像,因为这个模型能理解图像,你需要创建那些视觉提示、修改它们等等。碰巧我们就有这个东西,而且我们已经构建了好几年,它刚好可以用了。这是一个美妙的巧合,我们开始做更多这类 AI 演示。我们做了一个演示,你画画时,我们用实时图像生成器同步绘制。不知道你有没有看过。我们从中学到了很多。其中之一就是,这些东西很贵,我们不应该把所有的投资资金都花在对外演示上。
In order to take advantage of this wonderful technology, you're going to need a really good hackable canvas to work with images, because this model can understand images, and you're going to need to create those visual prompts, modify them, and all that stuff. It just so happened that we had the thing, and we'd been building it for years, and it was just ready to go. It was a wonderful coincidence, and we started doing more AI demos. We did one where you would be drawing, and we used a real-time image generator to draw in real time as you were drawing. I don't know if you ever saw that. We learned a lot from that. One was that stuff's expensive, and we shouldn't spend all of our investment money on demos out in the world.
所以你在生成式方面运气不错。有没有可能,就像我们现在认识到画布对于生成是显而易见的一样,我们也会把更多这类智能体式工作流看作“哦,那当然发生在画布上”?鉴于你可能不是完全靠运气,你有哪些方式在探索、推动或利用这个看似下一个大飞跃的机会?
So you got lucky on the generative stuff. What are the chances that, in the same way that we now recognize that a canvas is obvious for generation, we look at more of these agentic workflows as like, 'Oh, of course that happens on the canvas'? Given that you're probably not trying to 100% just get lucky, what are some of the ways you're trying to explore or push that forward or capitalize on what appears to be the next big jump?
我们在 2024 年初的对话是:我们认为这会走向何方?ChatGPT 已经出来了,图像模型也是。我想,好吧,每个人都觉得这将是定义我们这一代人的技术。我们在这个未来中的角色是什么?画布的角色是什么?ChatGPT 确实让人大开眼界,我们一直在思考它。我们想,聊天之所以有效,是因为聊天对人们有效。我可以和朋友、妻子、家人聊天。聊天对人类很有效,对 AI 也很有效。画布对人类很有效,但它也应该对 AI 有效吗?这是否会成为画布更受欢迎的原因,因为它已经擅长协作?它只是一个很好的协作环境。我们花了多年时间分析画布的独特和优点,但这让我们退后一步思考,为什么它实际上对协作有好处?我们在画布上能做哪些在聊天中做不到的事情?比如同时工作,但以一种你看不到别人在做什么的方式,但你知道它在那边,你可以过去看看,然后回来看看自己的东西。它正在发生,但不在你的视野或信息流中,不会分散你的注意力。你可以在同一文档中并行工作,这在 Microsoft 或 Google 表格中很难做到,但在画布上可以。你可以根据光标位置判断人们在做什么,也可以根据协作者的聚集看出谁在一起工作。你可以留下评论,让人们异步获取。我们想得越多,发现你总是可以聊天,可以轻松嵌入视频或音频,这些其他模态可以以比聊天更实时的方式接入。所以我们形成了这个论点:画布可能正是各种智能体一起工作的好地方。突然之间,你有了这样的东西:一张满是杂乱便签的截图,比如我的每周计划会议发生了什么?给我做一个结构大纲,把它放进数据库之类的。这在 2023、2024 年之前是不可能的。
The conversation we had at the beginning of 2024 was: where do we think this is going to go? ChatGPT had come out, the image models. I was like, okay, everyone's like, 'This is going to be the defining technology of our generation.' What's our role in this future? What's the role of the canvas? And ChatGPT was really just blowing people's minds, and we were all just thinking about it constantly. We were like, well, chat works because chat works for people. I can chat with my friends, with my wife, my family. Chat works really well for people. It just also works well for AI. The canvas works really well for people, but should it also just work for AI? Is that going to be a reason why the canvas becomes more popular, because it's already good at collaboration? It's just a good environment for collaboration. We spent years already kind of unpicking what is unique and good about the canvas, but it really made us step back and think about why it's actually good for collaboration. What can we do on the canvas that we can't do in a chat? Things like working simultaneously, but in a way that you don't see what other people are making, but you know it's over there, and you can come over and look at that, then come back and look at your own stuff. It's happening, but it's not in your view or your feed, it's not distracting. You kind of work in parallel even though you're working in the same document. It's very hard to do in a Microsoft or Google sheet, that type of parallel in the same document is very tough, but you could do it in a canvas. You could tell what people were doing based on where their cursors were. You could also see who's working together based on the clustering of collaborators. You could leave notes that people pick up asynchronously through comments. As we thought about it more, you could always chat. You could lay in video very easily or audio, these other modalities could just plug in in a way that makes more sense as a real-time thing than it would in a chat. So we developed this thesis around the canvas just might be a good place for intelligences to sort of work together. Suddenly you have something where you're like, here's a screenshot of a whole bunch of noisy sticky notes everywhere, like what happened during my weekly planning session? Make me an outline of that structure, put it into the database or something. That had never been possible until 2023, 2024.
所以我们意识到我们需要的是:我想在画布上与 AI 协作。我想要一些虚拟协作者。我希望画布上的一些光标是人,一些光标是 AI。我可能想要自己的私人 AI 助手,有些可能只属于画板或其他东西。它们应该能和我一起工作,看到我所看到的,做出我能做出的东西,但它们是 AI,就像我可以和 AI 聊天或在 Slack 中有一个 AI 成员一样。这就是愿景。我们要做到这一点。tldraw 将能够促成这件事,因为我们有所有的组件。今天可能吗?我们研究并原型化了,但今天还不可能。我们可以做原型,但机器人太差了。它们会和你下棋或跳棋,但会在错误的位置画 X,而且不知道。你问它们:“嘿,你画对位置了吗?”它们会说:“呃,没有,我画错了。我能注意到我画错了。让我画到正确的位置。”然后画到另一边去了。它们搞不懂坐标系之类的东西。我们试图限制和教导它们,想办法提示以获取正确信息,但它们就是很差。所以我们想:“好吧,它们做不到这个。我们不会发布一个有这样的功能。”如果你还记得,有一家公司叫 Diagram,Jordan Singer 的公司,最终卖给了 Figma,那是他们主张的一部分:“我们将拥有这些虚拟协作者。”他们很快遇到了同样的问题:“哦,它不像我们想的那样工作。”这不是模型擅长的。我们处于一个独特的位置,拥有一个相当低保真、有点傻气、有创意的画布。尽管这很糟糕,我们还是分享它。我们仍然构建了自动补全,开始发推文分享。我们尝试构建一些智能体式提示来在画布上生成内容,尽管它很糟糕。它会很烂,但也会很惊人,因为我们在内部都印象深刻,但它显然不是企业级软件。它远未达到我们可以真正支持并说“这和应用程序的其他部分一样成熟”的程度。
So the thing we realized we would need is: I want to collaborate with AIs on the canvas. I want little virtual collaborators. I want some of the cursors on the canvas to be people, and I want some of the cursors to be AIs. I might want to have my own private AI assistants, and some of them might just belong to the board or be other things. They should be able to work with me, see what I see, and make the same things I can make, but be AI, in the same way that I could chat with an AI or have a chat in my Slack with an AI member. That was the vision. We're going to do that. That's a thing that tldraw will be able to facilitate because we have all the bits and pieces there. Is it possible today? We looked into it and prototyped, and it is not possible today. We could prototype, but the bots were just bad. They would play chess or checkers with you, but they would draw X's in the wrong spot and wouldn't know it. You'd ask them, 'Hey, did you draw that in the right spot?' They're like, 'Uh, no, I drew that in the wrong spot. I can notice that I drew that in the wrong spot. Let me draw it in the right spot,' and it would be way over on the other side. They couldn't figure out the coordinate systems, all that stuff. We tried to limit it and teach it, figure out how to prompt in such a way that it got the right information, but they were just bad. So we're like, 'All right, so they can't do this. It's not like we're going to ship a feature that has this.' If you remember, there was a company called Diagram, Jordan Singer's company, ended up selling to Figma, and that was part of their proposition: 'We're going to have these virtual collaborators.' They ran into the same thing really quickly, which is, 'Oh, it just doesn't work the way we thought.' This is not something that models are great at. We were in this unique position of having a fairly low-fi, goofy, creative-looking canvas anyway. Even though this sucks, let's still share this. Let's still build autocomplete and start tweeting about it and sharing it. Let's try and build some of these agentic prompts to generate content on the canvas even though it sucks. It's going to be shitty, but it's also going to be amazing because we were all totally impressed by it internally, but it was clearly not enterprise software. It was clearly not anywhere near the maturity that we could really stand behind and say this is as much software as the rest of the app.
那真是又搞笑又有趣,还指向了未来。那个夏天我们做了自动补全功能。主要是 Lou Wilson 在做,Ryan Reed 也参与了。我们做了些视频,但效果太差没法发布,因为它会这样:我们告诉它“这是用户最后做的三件事”,让它预测接下来三件。你画个圆,它会说“好的,画了个圆,我不太清楚他在干嘛,让我继续看”。然后我在圆里面左边画个小圆,模型会说“我知道他在画脸,我要在右边再画个圆”。我按 Tab 接受,它就认为“哦,用户在画一串圆,我再画一个”。接着就像眼睛、眼睛、另一只眼睛,然后“哦,我现在真知道他在画圆了,继续吧”。它会画手臂、手掌、手指,然后觉得“哦,这像一棵树在生长,一个变五个,每个手指又该分出五个”。这有点像身体恐怖,但自动补全很棒。我们最终把那段代码——把画布转成文本,连同图像和摄像头信息发给模型——改编成一个应用:你有一个可以放在任何地方的框,带文本输入。你可以输入“画只猫”或“做个图表”,它就会在你定义的工作区内生成。它很安静。我们以 teach.taildraw.com 发布,你还能去玩。它变好了,因为底层模型进步了,就像 Make Real 一样。但这主要是演示“嘿,这是可能的”——让模型在 tldraw 里做那种“鹈鹕骑自行车”的绘画。很酷的是它生成的是和我能画的一样的基本图元。我常做的演示是让它画只猫,然后我在旁边画个高矩形,上面加个小黄矩形,说“让猫吹灭蜡烛”。模型看到的是截图和描述画布形状的 XML,没有坐标。但它会在猫嘴处画蓝色线条,删除黄矩形,画些烟雾,然后说“我做到了,猫吹灭了蜡烛”。
It was just hilarious and entertaining, and it pointed toward the future. We did autocomplete that summer. Lou Wilson mainly worked on it, Ryan Reed also worked on it. We made some videos, but it was too bad to ship because it would do things like: we'd tell it 'here are the last three things the user did' and ask for the next three. You draw a circle, and it would say 'alright, drew a circle, I don't really know what he's doing, let me keep watching.' Then I'd draw a smaller circle inside, off to the left, and the model would say 'I know what he's doing, he's drawing a face. I'm going to draw another circle off to the right.' I'd hit tab to accept, and that would be sent to the model. Then the model would think 'oh, the user is drawing a line of circles, let me draw another one.' It would be like eye, eye, another eye, and then 'oh, now I really know he's drawing circles, let's just go.' It would draw an arm, then this part, then the hand, then fingers, and it would think 'oh, it's like a tree being expressed, where one becomes five, now each of these should become five like fingers on fingers.' It was like body horror, but autocomplete was awesome. We eventually took that code—turning the canvas into text and sending it to the model along with the image and camera info—and adapted it to an app where you had a box you could place anywhere, with a text input. You could type 'draw me a cat' or 'make me a diagram' and it would generate within that work area. It was very quiet. We launched it as teach.taildraw.com. You can still go there and play with it. It's gotten better because the models have improved, just like Make Real has. But it was mostly a demo of 'hey, this is possible'—getting a model to do the kind of 'pelican riding a bicycle' drawing thing in tldraw. It was cool that it generated the same primitives I can make. The demo I always do is have it draw a cat, then I draw a tall rectangle next to it with a little yellow rectangle on top and say 'make the cat blow out the candle.' The model sees a screenshot and XML describing shapes, not coordinates. But it will make little blue lines coming from the cat's mouth, delete the yellow rectangle, draw some smoke, and say 'I did it, the cat's blowing out the candle.'
这非常糟糕但又神奇。没人会因为这东西会用画布而失业。猫画得差,烟雾也差,甚至算不上好,但你会觉得“哇,这里有点意思”。我思考了很多如何以符合这些 AI 技能水平的方式来叙事。如果我说“这些是你的虚拟协作者”,它们并不好。它们会做人类不会做的疯狂事。所以把它们称为虚拟协作者似乎不对。我想“也许它们是你要召唤的鬼魂或精灵”。我有些思考时画的草图。然后我想“它们可以是小虫子或小精灵。它们很小,适合光标大小。它们不是人,但有点人形。它们不绑定于你”。结果发现有很多种精灵。斯堪的纳维亚精灵还行但不友好。爱尔兰精灵很可怕——别去了解,它们会偷孩子、从烟囱下来。英国精灵很迷人,像迪士尼传说中的小叮当。所以我想“它们肯定是英国精灵。这很合适,我们在伦敦”。它们有能力和配饰。随着深入,处理智能体的问题让我明白为什么画布是协作的好地方。用五个智能体进行氛围编程的一个大问题是记不清哪个智能体在做什么、有什么上下文。区分它们很难——满屏都是运行中的聊天,你会想“那个在做认证,这个在设置项目描述”。这个问题在画布上很好解决,因为它们看起来不同。给它们不同的帽子、翅膀图案、衣服和颜色。
It's extremely shitty but amazing. No one's losing their job over this thing knowing how to use the canvas. The cat is a bad drawing, the smoke is bad, it's not even good, but it's like 'wow, there's something really interesting here.' I thought a lot about how to frame this narratively appropriate to the skill level of these AIs. If I said 'these are your virtual collaborators,' they're just not good. They do crazy stuff no human would do. So framing them as virtual collaborators seems wrong. I thought 'maybe they're ghosts you're invoking, or spirits.' I have little drawings I sketched while thinking about this. Then I thought 'they could be little bugs or fairies. They're small, they fit the cursor size. They're not people, but kind of humanoid. They're not bonded to you.' Turns out there are many types of fairies. Scandinavian fairies are okay but not very friendly. Irish fairies are terrifying—don't read about them, they steal your kids, come down your chimney. English fairies are charming, like Tinkerbell from Disney lore. So I thought 'they're definitely English fairies. That's appropriate, we're in London.' They have powers and accessories. As we went deeper, dealing with agents taught me why the canvas is such a good place for collaboration. One big problem with vibe coding with five agents is remembering which agent is doing what and what context they have. Telling them apart is hard—you have a screen full of running chats, and you're like 'that one's doing authentication, this one's setting up the project description.' That problem works well in the canvas because they just look different. Give them different hats, wing patterns, clothes, colors.
还有就是我所有智能体的状态——哪怕只有一个——它是在等我?还是在思考?在做什么?在工作吗?我希望所有智能体都全力运转,一刻不停。如果我真想推进事情,我不希望任何人在等待。系统应该让这变得容易。所以这是我们可以可视化呈现的另一部分状态。比如它们在做什么?是在写东西还是在做线框图?画布让这变得非常简单,原因和它对人类友好一样。你直接就能看到小精灵在哪儿,光标在哪儿。
There's the what is the state of all of my agents that I'm trying to run at the same time, even if it's just one. Like, is it waiting for me? Is it thinking? Is it doing something? Is it working? I want all of my agents to be totally maxed out and running at all times, right? I don't want anyone to be waiting if I'm really trying to push things. And the system should make that easy. So that's another part of the state that we can represent visually. There's like what are they working on, you know, like are they writing this thing or are they making wireframes? The canvas makes it very easy for the same reason why it's good for people. You just see where the fairy is, see where the little cursor is.
同时向多个智能体分配任务或编排这类系统的复杂性——这正是我们最近在做的事。我在里斯本演示的就是这个。我想给一个任务,它超出了任何一个智能体单独能处理的范围,然后我希望这些智能体——这些小精灵——能围绕解决这个任务自我组织。比如我说:“嘿,我想让你为我的应用创建线框图,但还需要你根据这些输入写 PRD,然后还要你做一个图表,展示这些线框图如何交互,比如用户流程。”它们就会开始扇动翅膀表示正在工作。它会思考,画布上的小角色会摸摸下巴,敲敲下巴,然后说:“嗯,这活太多了,我召唤其他小精灵来帮忙。”然后你对话的那个小精灵就会切换成编排模式,开始创建任务并分配给系统中的其他智能体去执行,同时还会检查所有任务的进度,在别人完成后分配更多任务。实现这些的规则相对简单,但突然之间,画布上就出现了这种疯狂的涌现行为。我演讲时用了八个智能体同时工作。虽然有点混乱,但太棒了。这远远超出了我在其他范式下做过的任何事情。不好的地方在于,模型不知道文字有多大,也不知道不能可靠地把东西叠放在一起,或者把内容分散开。好的地方是,我知道每个人在做什么,我能区分它们,也能知道什么时候出了问题,然后说:“哦,不不不,停,别建了,那不对。坏精灵。”诸如此类。所以这个论点并不新颖,就是:对人类好的东西,对 AI 也会好。
The complication of addressing tasks to multiple agents at the same time or orchestrating these types of systems. This is what we're doing very recently. This is what I was just demoing here in Lisbon. I want to give a task that is more than any one of these agents should work on at the same time. And then I want the agents, the fairies, to self-organize around solving that task. Say, 'Hey, I want you to create wireframes for my app, but also I need you to write the PRD based on these inputs, and then I want you to make a chart of how those wireframes interact with each other like a user flow.' And they'll start flapping to represent it working. It'll think, it'll kind of do one of these—a little character on the canvas touches his chin, taps his chin and be like, 'Well, that's too much work for me. Let me summon some other fairies to assist me.' And then that fairy you talked to will kind of switch into an orchestrator mode where it is now creating tasks and assigning tasks to the other agents in the system to go do those things, and also checking in on the progress of all those things and assigning more tasks when people finish. The rules to enable that are relatively small, but suddenly this crazy emergent behavior happens on the canvas. The talk that I gave, I had like eight agents all working at the same time. It was kind of chaos, but it was awesome. It was well beyond what I'd ever done on any other paradigm. The parts that were bad were still like, yeah, the model doesn't know how big text is or doesn't know not to put things on top of other things reliably and to spread stuff out. The parts that were good were that I knew what everyone was doing. I could tell the difference between all of them. And I could know when things were going wrong also and be able to say, 'Oh, no, no, no. Stop, stop building that. That's just wrong. Bad fairy.' All those types of things. So, the thesis isn't very creative. It's just like, well, it's good for people. It'll be good for AI too.
再次看到你从第一性原理出发思考每一个细节,甚至包括小精灵的姿势作为它们正在做什么的信号,这真的很酷。在那个更智能体化的基础里,有太多微决策了。听你讲你的思考过程,我很欣赏。
And it's cool to see again, you just think through every little detail from first principles, even down to the posture of the fairies as a signal for what they're doing. Like, there's so many micro decisions in that more agentic foundation. I'm appreciating listening to you talk about just your thought process even for it.
我很幸运能和一群非常有创意的人一起工作——主要是 Max、Drake、Lou Wilson,还有 Mima Kabalo。我围绕自己组建了这个团队,所以不光是 Stephen 的想法。我们在设计这些功能时,会有一些最疯狂的对话。比如,如果画布上有个池塘呢?一个魔法池塘?然后大家就说,对对对,因为你会想要一个小精灵来管理池塘。如果有东西进入池塘,它就应该处理那个东西。那应该成为这个智能体的提示。换句话说,建立一个领域、一个文件夹之类的,然后让一个 AI 智能体来管理那个文件夹的内容。如果有东西进入这个文件夹,我希望你根据我们商定的规则运行这个工作流。这完全说得通。
I mean, I'm very lucky to work with some other really creative people—Max, Drake, Lou Wilson mainly on this I'll say, so Mima Kabalo. I've kind of built that team around me, so it's not just Stephen's ideas. But we do have some of the craziest conversations when working through these features. Like, all right, what if there was like a pond on the canvas, you know, like an enchanted pond? And it's like, yeah, yeah, yeah, because you'd want to have a fairy that kind of runs the pond. And if anything enters the pond, it should operate on that. That should be the prompt to this agent. In other words, establishing a domain, a folder, whatever, and having an AI agent that essentially manages the contents of that folder. If anything enters this folder, I want you to run this workflow based on what we've agreed. Makes total sense.
好的,没错。就像,嗯,那是个池塘,魔法池塘——不可能是别的,只能是魔法池塘。
Okay. Yeah. Like, yeah, it's a pond, you know, enchanted—can't be anything else than enchanted pond.
隐喻实际上是双向起作用的。有些想法从技术层面出发,从 AI 世界出发:好,我们怎么实现 MCP?比如,也许小精灵会扭曲进入 Notion,就像它的眼睛翻白,因为它正在访问信息,无法做其他事。也许它向一只蝴蝶低语,蝴蝶就飞走了。我们需要某种隐喻来表示它从用户无法直接访问的另一个来源获取信息。然后有些想法则从另一个方向来:比如,小精灵有魔杖,我们能拿它做什么?我在读关于精灵民间传说的 PDF,发现很多精灵会给人类留下礼物。精灵传说里有很多送礼和赠礼的情节。我们怎么实现呢?然后就想,哦,你可以给小精灵留东西。留什么呢?比如额外的上下文。如果你在这个区域工作,请先读这个——这在智能体编程中很常见:小上下文文件、智能体文件,或者“在修改我的测试之前请先阅读我”之类的文件。然后就想,对,我们可以在画布上实现这个——小卷轴、小纸条、小信件之类的——它们可以互相留东西。这真是个很棒的主意。所以拥有这层隐喻实际上对产生想法非常有用。对我们来说就是这样。
A metaphor actually works in both directions. Some things will start from the technical side, the AI world: okay, how are we going to do MCP? Like, all right, well maybe the fairy is kind of warping into Notion, warping, you know, like its eyes roll back in its head as it's accessing this information because it's not going to be able to do anything else. Maybe it whispers to a butterfly and the butterfly flies. We need some sort of metaphor for this accessing information from a different source that's not accessible to the user directly. And then some of it goes in the other direction: saying like, yeah, fairies have wands, you know, what do we do? Can we do anything with that? Like, I'm reading PDFs about fairy folktales and I'm like, yeah, a lot of fairies leave gifts for people. There's a lot of gift leaving and gift giving in fairy lore. How would we do that? And be like, oh, you know, you could kind of leave things for the fairy. Yeah, what would you leave? Oh, you know, maybe you'd say like extra context. If you're working in this area, make sure that you read this first, which is something that we do all the time in agentic coding: little context files or agents files or 'please read me before working on my tests' type of files. You're like, yeah, let's just do that on the canvas—little scrolls or little notes or little letters or something like that—and they could leave them for each other. And it was like a great idea. So having that layer of metaphor is actually really great for coming up with ideas. It has been for us.
我记得和 Smith and Addiction 的 Mike Smith 聊过,他说一旦你找到了那个隐喻,它就会开始展开,加速构思,因为它会不断叠加。我能看出你把自己放在了正中间。你给自己建立了一个很好的精灵传说基础,这本身就是你作为建造者的一个缩影。你深入研究了精灵传说,但现在你已经到了想法快到跟不上的地步。
I remember talking to Mike Smith from Smith and Addiction and he talked about how once you hit on the metaphor, it just starts unraveling and it accelerates ideation because it starts to compound. And I can tell you put yourself right in the middle. You know, you gave yourself a nice little fairy lore foundation, which in itself I think is a microcosm of who you are as a builder. The fact that you've gone so deep into fairy lore, but you're at the point now where it's like, oh, it just you can't even keep up with how fast the ideas are coming.
太棒了。我们的线框图艺术、关于这些东西的文档。我们在构建这些东西时会在 Twitter 上发很多内容。Max 和 Lou 也是这方面值得关注的人,因为我们都是边做边摸索。
It's great. Our wireframed art, our tro docs on this stuff. We post a lot on Twitter as we're kind of building these things. Max and Lou also like great follows for this stuff because we're just figuring it out as we go.
但这其中涉及太多东西了——真的很难忽视,我们正在全力摸索。你知道,我们必须提醒自己,我们正在解决的问题是:如何让人类 + AI + 多个 AI + 多个人类在单个文档中实时协作。很少有公司或产品在攻克这个问题;我们处于一片巨大的未知领域。但同时,一些最雄心勃勃的事情正在软件领域发生,对吧?所以如果你想解决一个大问题,我们就这样一头扎进去了。
But there's just so much—it's really hard to miss, and we are absolutely figuring this out. You know, we have to remind ourselves that the thing we're figuring out is how to do human plus AI plus multiple AIs and multiple humans collaborating in a single document in real time. There are very few companies or products working on that problem; we are in massively uncharted territory. But also, some of the most ambitious stuff is happening in software, right? So if you wanted to tackle a big problem, we just kind of launched ourselves into it.
是的。
Yeah.
有一些非常有趣的问题,比如如何处理编排?如何处理等待、轮流发言,或者跟踪 AI 会处于的不同状态?这实际上非常类似于视频游戏 AI,比如《星际争霸》或 CRPG 这类策略游戏,我们参考了很多,它们已经解决了很多这些问题。你有这些小实体,它们有自己的状态机:我在等待、我在行动、我在移动、我在巡逻,诸如此类。它们根据敌人接近在这些状态之间切换——现在它们会进入攻击状态,现在它们会切换到战斗模式之类的。我们也在做同样的事情。只不过我们现在还使用 AI 来控制它们在该状态下的行为。但这仍然涉及大量程序化的状态切换,以及为这些“精灵”设置角色和模式,我认为这也会在其他地方出现。我的意思是,这似乎是解决多智能体或编排问题的唯一方法。在代码中也是一样。只是从某种程度上说,这几乎不可能看到,不可能管理,或者至少非常困难。希望他们能解决这个问题,因为看到它们一起工作真的很酷,我可以想象拥有一群能够互相交谈、互相发邮件的 AI 程序员,这将是值得的。
There are really interesting problems like how do you handle orchestration? How do you handle waiting, turn-taking, for example, or following the different kinds of states that the AI is going to be in? It's actually very similar to video game AI, where you have these little strategy games like Starcraft or CRPGs or something that we look a lot at, which have solved a lot of these problems. You have these little entities that have their own little state machine: I'm waiting, I'm acting, I'm moving, I'm patrolling, all these things. And they kind of move between those states based on an enemy approaching—now they'll aggro, now they'll switch into a fighting mode or something like that. And we're doing the same thing. It's just that we're also using AI now to control the behavior while they're in that state. But it's still a lot of programmatic switching between things and setting up those roles and modes for these fairies, which I think is something you're going to see elsewhere. I mean, it just seems to be the way to solve this problem—the multi-agent or orchestration problem. It would be the same in code. It just would be impossible to see in a way, impossible to manage, or at least very difficult. Hopefully they figure it out, because it's really cool to see them working together, and I can imagine having a fleet of AI coders who can talk to each other and email each other or whatever is going to be worth it as well.
是的。
Yeah.
同样。事情变得清晰了,我很感激你让我们稍微了解你的想法。我觉得在过去一个多小时里,我有点了解你和你的思维方式了,我非常享受。所以,感谢你的时间,我期待看到更多 Twitter 演示,因为它们绝对是我在那个小鸟应用上最喜欢的部分。现在,在你离开之前,我想花一分钟向你介绍我最喜欢的产品,因为我经常被问到我的技术栈。Framer 是我构建网站的工具。Genway 是我做研究的方式。Granola 是我在评审期间做笔记的工具。Jitter 是我为设计制作动画的工具。Lovable 是我用代码实现想法的工具。Mobin 是我寻找设计灵感的地方。Paper 是我像创意人一样设计的工具。而 Raycast 是我每一步的快捷方式。我精心挑选了这些公司,这样我就能全职做这些节目了。所以,支持这个节目的最佳方式就是去看看它们。你可以在 dive.comclub/partners 找到完整列表。
Same. It's becoming clear, and I just appreciate you letting us get in your brain a little bit. Like I feel like I got to know you and the way that you think and process things a little bit over the last hour plus, and I've thoroughly enjoyed it. So, I appreciate your time, and I will look forward to more Twitter demos, because they're definitely my favorite parts of that little bird app. Right now, before I let you go, I want to take just one minute to run you through my favorite products because I'm constantly asked what's in my stack. Framer is how I build websites. Genway is how I do research. Granola is how I take notes during crit. Jitter is how I animate my designs. Lovable is how I build my ideas in code. Mobin is how I find design inspiration. Paper is how I design like a creative. And Raycast is my shortcut every step of the way. Now, I've hand selected these companies so that I can do these episodes full-time. So, by far the number one way to support the show is to check them out. You can find the full list at dive.comclub/partners.