Designing Neural Interfaces: From Apple Watch to Brain-Computer Interaction
打开互动全文版(中英对照 + 朗读 + 问答)→Neuralink 设计工程师 Ruse Madavian 分享他从设计 Apple Watch Siri 表盘到构建脑机接口的历程,探讨前沿界面如何需要从第一性原理重新思考交互。
Ruse Madavian, design engineer at Neuralink, shares his journey from designing the Siri watch face at Apple to building brain-computer interfaces, exploring how frontier interfaces require rethinking interaction from first principles.
我们在 B2B SaaS 中经常谈论工艺,但设计前沿界面是怎样的体验?
We talk a lot about craft in B2B SAS, but what's it like designing frontier interfaces?
我认为这个世界将出现新的交互模式,它们与我们必须用物理手来操作某物截然不同。
I think there will be new interaction models in this world that are yeah, just very different than having to use our like physical hands to like articulate something.
脑机交互的未来如何塑造设计实践?
How does the future of brain computer interaction shape the practice of design?
我们今天称之为“氛围编程”。你知道,在未来,模型实际生成你所描述的产物所需的时间,我认为会缩短到一帧。这将是一种非常激动人心且狂野的体验:你只需坐在电脑前,基本上与它一起做白日梦,屏幕上就会弹出直接延伸你脑海想法的东西。
What we today call vibe coding. You know, in the future, like the time that it takes for a model to actually produce an artifact that you described will, I think, come down to a frame. This will be a very exciting and wild experience where you can just sit down on a computer and basically daydream with it and things will pop up on screen that are a direct sort of extension of what you have in your head.
欢迎收听 Dive Club。我是 Rid,这里是设计师永不停歇学习的地方。本周的嘉宾是 Neuralink 的设计工程师 Rooz Mahdavian。所以,我们将深入探讨设计神经接口的体验。比如,如何让某人仅用思维控制光标?这一集我最喜欢的部分之一是听到 Rooz 如何从第一性原理思考人机交互的每一个细节。设计思维的水平令人印象深刻。但在深入之前,让我们先听听 Rooz 的经历,因为实际上是从 Apple 开始的。
Welcome to Dive Club. My name is Rid and this is where designers never stop learning. This week's episode is with Ruse Madavian who's the design engineer at Neurolink. So, we're about to get pretty nerdy and talk about what it's like designing neural interfaces. like how do you allow someone to control a cursor using just their mind? And one of my favorite parts of the episode is hearing how Ruse has had to think through every single detail of human computer interaction from first principles. Like the level of design thinking is really impressive. But before we get into all of that, let's hear a little bit more about Ruse's journey because it actually starts at Apple.
我认为这段旅程最早可以追溯到高中一年级。当时伯克利有一项研究,真的非常疯狂。他们基本上让人们进入 fMRI 机器,让他们观看 YouTube 视频,收集了数小时的数据,然后进行某种所谓的机器学习,利用观看视频时产生的 fMRI 信号,预测他们正在观看的电影帧,并实时将这些帧重建成看起来像他们所想内容的电影。结果非常狂野,非常赛博朋克。它是由帧组成的画面,从中可以看到语义性的东西。比如,你会看到一只鸟,或者一片海滩,从噪声中浮现出一些形状。然后他们又做了一个更疯狂的步骤:让这些人在 fMRI 舱里小睡,然后你实际上可以看到他们在做什么梦。这让我一旦看到就难以释怀,简直难以置信。在那之前,我花了很多时间制作电影。所以,能够直接向他人展示你脑海中的图像,对我来说非常神奇。因为制作短片时,你要花很多时间试图重现脑海中的画面。而能够直接展示给别人,简直不可思议。我想那是我第一次接触到某种神经接口的概念。另外,我也认为计算机本身就是一个神奇的东西。所以我深入研究了那个领域。那是另一种将脑海中的东西展示给别人、为他人创造体验的方式。
I think the earliest point in that journey for me was uh like freshman year of high school. There was this like research that had come out of Berkeley that was truly wild. Uh they had basically put folks through an fMRI machine and they had them I think it was watch YouTube videos uh and they just collected hours and hours of this data and then they ended up doing some form of like quote unquote machine learning where they could take this fMRI signal that would happen while they were watching these movies and then predict what frame of the movie that they were watching and then in like real time they could then reconstruct all of these frames into what to our eye looks like then a movie of what they're thinking. The results were super wild, like super cyberpunk. It was like this composition of frames from which you could see these like semantic things come out. So you could see like, oh, there would be like a bird or it would be like uh, you know, a beach, like you could see some shapes emerge from this noise. And then they took this extra crazy step of like them having them go like take a nap inside of one of these fMRI pods and then you could see in effect like what they were dreaming. This was like something that once I saw it was like hard to like let go of. It was just like unbelievable. Up to that point, I'd spent a lot of time actually making movies. So, the idea that you have these images in your mind that you can now show other people directly was like really really magical to me for that reason. It's like you spend all this time trying to recreate that when you go through the process of like making a short film. The idea that I could just show somebody that directly was just, you know, incredible. That was like I guess my first exposure to the idea of some form of a of a neural interface. And then separately, you know, I was also I think computers in general are a magical thing. So I was I was really going down that rabbit hole. That's another way in which you can take something that is in your mind and then show it to somebody else and create an experience for somebody else.
快进 5 年。我在大学期间有一个疯狂的机会,基本上是在 Apple Watch 团队实习。表盘团队,那是在手表发布一年后。所以那真的是手表界面早期探索的阶段。表盘团队也是一个非常棒的地方,因为那是你一直能看到的东西,而不是隐藏在界面层之下。我的实习项目后来成为了 Siri 表盘。我当时非常兴奋,想象一个主动式计算机是什么感觉,一个始终陪伴你的东西。我们有一个概念叫“手表上的路过”,就是你查看时间——手表的基本价值主张就是查看时间——但现在有机会在查看时间的同时呈现其他信息。所以计算机在某种程度上不再是你在主动使用的东西,而是淡入背景,希望在你外出时为你做有用的事情。第一个版本的 Siri 表盘正是针对这一点,它使用了比我们今天更原始的机器学习形式,但试图找出我们能给你的最相关信息。无论是即将到来的日历事件这种老套的例子,还是更情境化的东西,比如如果你一直在散步,我们可以显示开始锻炼的选项。在 Apple 这样的公司经历整个流程非常有趣,从脑海中的概念,到初步探索,到实际原型构建,再到迭代循环,最终发布一个很棒的产品。所以我回来又待了一年进一步开发,毕业后全职加入继续这项工作。
So fast forward 5 years. Uh I had this crazy opportunity in college to uh basically intern on the Apple Watch team. The faces team, this was like one year after the watch had come out. So it was really the early days of what the interface for watch might look like. And the faces team is also a super amazing place to be. That is like the thing you actually see all of the time as opposed to being sort of more buried into the layers of the interface. My intern project was what would go on to become the Siri watch face. So I was just yeah generally really excited about what a proactive computer might feel like something that is with you all the time. We have this notion of driveby on watch which was you're checking the time that is in some ways the fundamental value proposition of a watch is that you go to check the time but now there's this like kind of opportunity to surface something else while you're checking the time. So the computer in a way is not something you're actively using anymore, but something that sort of fades into the background uh and can hopefully just do useful things for you while you're out and about in the real world. The first version of the Siri watch face was sort of aimed at exactly that, which is like again more primitive forms of machine learning than uh I think we have today, but uh a way for us to try and figure out what is the most relevant piece of information uh that we can give you. Whether it is like a calendar event that you have coming up the sort of cheesy examples but also more contextual things like you know if you have been going for a walk like we can surface the ability to start a workout right that was just a ton of fun to go through that whole uh flow at a company like Apple I guess like my first experience going end to end from just a concept that you have in your head to the initial explorations of what that might feel like to the actual prototyping work building it and then living on that and you sort of go through this loop over and over again ultimately to ship something that feels awesome. So I came back for another year to build that out further and then I joined full-time after graduating to just sort of continue that work.
在那里待了一年后,我仍然在思考高中时最初接触神经接口的经历,那肯定还在我脑海深处。当然,Neuralink 在几年前就已经公开宣布,并且仍在进行大量工作。所以在他们第一次演示之后,那是 2019 年 7 月。观看之后,很明显这将成为人们实际使用的真实事物。大约在 2019 年底,我从手表上的一个前沿界面跳到了另一个界面。
So after a year there I was still thinking about that initial sort of exposure to neural interfaces in high school was definitely like in the back of my mind still and of course Neuralink had publicly announced a couple years prior and was still doing a lot of this work. So after their first demo, this was like July 2019. It was just watching that it was clear that this is actually going to be a real thing that people get to use. It was like late 2019 that I made the jump from one Frontier interface uh on watch to another in this case interface.
快速插播一条消息,然后我们继续。前几天我看到一条让人停下滑动的推文。Tailwind 的创建者直接在 Paper 上工作,以训练输出达到完美。他们甚至投资了这家公司。所以想想可能性。未来,你可以在 Paper 中设计东西,然后右键复制完美的 Tailwind,就像创建者亲手编写的一样。或者,你可以将现有的代码组件导入 Paper,直接在画布上进行编辑。我的意思是,这将彻底改变我们设计和交付 Web UI 的方式。
Real quick message and then we can jump back into it. I saw a scroll stopping tweet the other day. The creators of Tailwind are working directly on paper to train the output to be perfect. They even invested in the company. So just think about the possibilities for a second. In the future, you could design something in paper and then just rightclick and copy the perfect Tailwind as if the creators themselves wrote it by hand. Or maybe you take an existing code component and import it into paper to make edits directly on the canvas. I mean, this is going to totally change how we design and deliver UIs for the web.
这也是我为什么大力押注 Paper 作为下一个伟大设计工具的又一个原因。你今天就可以试试看,只需访问 dive.club/paper。大家都知道,Jitter 多年来一直是我做动画的首选工具,但他们仍然在疯狂地推出新功能。我是说,就在这个夏天,他们发布了评论、钢笔工具、变形、文本渐变、Google 字体等等。所以,如果你还没试过,我保证你会震惊于在 Jitter 中将设计变为现实是多么容易。今天就试试吧,前往 dive.comclub/jitter。过去 15 年我每天都在设计产品,但在过去 6 个月里,一切都变了。有了 AI 的加入,我比以往任何时候都更快地产生想法。但如果我无法获得让团队对齐所需的反馈,这一切都无关紧要。而目前,获取异步反馈仍然很糟糕。所以,我正在构建我一直想要的产品,它叫 Inflight。我每天都用它来分享想法并从团队获得反馈,它彻底改变了我的工作方式。我很兴奋能展示给你看。现在,我只对 diveclub 听众开放访问权限。所以,前往 dive.comclub/inflight 抢占你的名额。好了,现在进入正题。
And it's just another reason why I'm betting big on paper as the next great design tool. You can try it out today. Just head to dive.club/paper. By now, you know that Jitter has been my go-to tool for animation for years now, but they are still shipping like crazy. I mean, just this summer, they've released comments, pen tool, morphing, text gradients, Google fonts, and a bunch more. So, if you haven't yet, I promise you will be shocked at just how easy it is to bring your designs to life in Jitter. So, go ahead and give it a try today. Just head to dive.comclub/jitter. I've been designing products every day for the last 15 years. But in the last 6 months, everything has changed. With AI in the mix, I'm cranking out ideas faster than ever. But none of that matters if I can't get the feedback that I need to get the team aligned. And right now, getting async feedback still kind of sucks. So, I'm building the product I've always wanted, and it's called Inflight. I use it every day to share ideas and get feedback from the team. and it's totally changing the way that I work. So, I'm excited to show you. Right now, I'm only giving access to diveclub listeners. So, head to dive.comclub/inflight to claim your spot. Okay, now on to the episode.
我很喜欢“前沿界面”这个词,因为在 LinkedIn 上作为一个条目,你知道,我们有过相似的职位头衔,但你却在从事完全不同类型的产品和用户体验,从头到尾每一个部分都是独一无二的。
I love the phrase frontier interface because on LinkedIn as a line item, you know, we've had similar job titles and yet you're working on fundamentally different types of products and user experiences and like every piece top to bottom is unique.
我认为,所谓的前沿界面,就是你完全不知道实际的交互模型会是什么样子。历史上,所谓的“计算机”的交互模型最终定义了你能用它做的一切。回顾过去 60 年,那个交互模型通常只是输入机制的函数。你知道,是最初的光笔,只是屏幕上的一个点?还是光标,加上点击能力以及一定数量的点击?还是像 iPhone 那样的直接触摸?也就是说,你所有 10 根手指同时追踪所有 10 个点,这最终定义了计算机能做什么的形状,因为你如何表达你的意图、以什么节奏和分辨率表达你的意图,将决定你能用计算机做什么。在很多方面,我们可以超越物理层面。我们可以捕捉这种意图,现在我们脑海中有一个图像。我们现在如何重新创建它?我们把它分解成成千上万个运动意图,我们移动手来有效地给屏幕上的一个像素上色,在某种抽象层面上,我们非常手动地经历这个过程,直到最终我们能在计算机上看到最初脑海中的某种形式。如果你有一个神经接口,你就不必那样做,就像看向远方一样,因为那个意图在堆栈中比我的手高得多。你可以直接读取它。这就是我所说的前沿。我们真的不知道交互模型可能是什么样子,但设备本身有可能改变那个模型。所以,这仅仅是探索整个想法空间。
I think what that really means to be like a frontier interface is one that you have no idea what the actual interaction model is going to look like. Historically, like the interaction model for a quote unquote computer is what ends up defining all the things that you can do with it. Looking back like 60 years, that interaction model is usually a function of just what the input mechanism is. And you know, is it the light pen originally like just some point on the screen? Is it the cursor which is like you know that plus the ability to click and some number of like clicks or is it direct touch like the iPhone? So like you know all 10 of your fingers at the same time tracking like all 10 of those points that ends up really defining the shape of what the computer can do because how you express your intent and at what cadence and like resolution you can express your intent is going to define like what you can do with a computer. In many ways we can go you know potentially beyond just the physical. We can capture this intent that right now we have this image in our mind. How do we go about recreating that right now? We like decompose that into thousands, tens of thousands of, you know, motor intents where we move our hands to effectively color a pixel on the screen at some level of abstraction and we go through that process very manually up until finally we can then see some form of what originally was in our minds on the computer. If you have a neural interface, you don't have to do that necessarily, just like looking far down the horizon because that intent is there much further up the stack than my hand is. You could potentially read that directly. So that's what I mean when I say frontiers. We really don't know what the interaction model could look like, but the the sort of device itself can potentially change that model. So it's just about exploring that whole space of ideas.
那么快进一点。你是在什么时候意识到有机会进入一个更传统的设计角色?
So fast forward a little bit then. What was the point where you realized that there was an opportunity for you to step into more of a traditional design role?
我开始做了大量的原型设计,来感受下载一个像 Neuralink 这样的应用、连接到你的植入物并首次校准模型会是什么感觉。所以,最早的一步就是感受一下人类体验可能是什么样子,以及这是否是我们都非常兴奋的北极星。我花了很多时间纯粹在 iOS 上做这件事,因为那是当时的重点。感觉又像回到了苹果的早期,你有一块真正的空白画布,然后你一遍又一遍地循环:最初脑海中的一些概念,然后你知道的 10 个不同的迭代,最终变成你可以拿在手里把玩的东西。那真的很酷。但另一方面,因为过程太早期,几乎没有足够的约束来真正做出有意义的东西。你可以做很多很酷的事情,但最终你对交互模型的猜测有点太多了。显然,这不是为我们自己设计的。我们是为第一个参与者设计的,但那个人会有某种运动障碍。所以,他们的感官体验会和我们非常不同。因此,如果你无法感受到它,就很难真正直接地共情并为此优化。
I started doing just a ton of prototyping on what it would feel like to, you know, download an app like a Neuralink app and connect to your implant and sort of calibrate a model for the first time. So like the the earliest steps in here was just like getting some feel for yeah what a human experience could be and if this is like a northstar that we're all really really stoked about. Spent a lot of time doing that purely on iOS cuz that was the focus at the time. It felt like those early days at Apple again where it was like you have just a truly a blank canvas and then you're just going through the loop over and over again of like some concept that you have initially in your mind and then you know 10 different iterations on that thing and then ultimately something that you can hold in your hand and play with. So that was really cool. The flip side is though that because it was so early in the process there were almost like too few constraints to actually do something really meaningful here. It was like you could do a lot of cool things but ult like you're guessing a little bit too much on what the interaction model is going to be. And obviously it is something that we're not designing for ourselves. We're designing for, you know, whoever that first participant is going to be. But it's going to be somebody with some sort of motor disability. So it's like they're going to have a very different sensory experience than we do. So it's really hard to actually try and directly empathize with that and actually optimize for that if you can't feel it.
我能深入探讨这一点吗?因为这是你说话时我想到的,你在做所有这些不同的探索。是的,也许有点早,但让我着迷的部分是,我们作为设计师经常谈论共情,但这完全是另一个层次,你知道吗?非常困难。你不仅仅是在想象一个你从未有过的工作,你是在想象一种你从未接近过的存在状态。那是什么感觉?
Can I drill in on that point for a second? Because that was something that was coming to mind while you were talking where you're doing all these different explorations. Yeah, maybe it's a little bit early, but the part that's so fascinating to me is like we talk a lot about empathy as designers, but this is a whole other level, you know? It's like very difficult. You're not just imagining a a job that you've never had before. You're imagining a state of being that you've never really come close to before. What's that like?
我甚至想说,这不是你能想象得足够好以至于有用的东西。我认为我们有自己的内部模型,关于基本上只是尝试想象不动是什么感觉。这在某种意义上很有效,因为你可以只是坐在那里,比如对于校准任务,当你实际上只是试图想象移动时。你可以做到这一步,而且我认为你可以做得很好,因为你不需要移动。所以你只是视觉上尝试专注于屏幕上的某个东西,感受一下当你在屏幕上观看某物时,你的脑海中是否有什么事情发生,感觉几乎像镜像神经元——我松散地使用这个词——但镜像神经元在放电,你感觉到你和那个东西之间有联系。那一步你肯定可以想象,因为你可以体验到,这就是我们所说的开环。这意味着,如果你想象一个卡通循环:里面有一个人,那个人脑海中有一个意图,那个意图变成某个动作,然后他们观察那个动作的结果,然后结果回到他们的大脑。所以,就像有一个循环控制。
I would go so far as to say that it is not something that you can imagine well enough for that to be useful. We have our own I think internal models of like what it would feel like to yeah basically just try and imagine something without moving. That works well in the sense that you can just sit there obviously for something like a calibration task when you're actually just trying to imagine moving. You can do that step of it and I think you can do that step of it well because you don't have to move. So you just kind of visually try and focus on something on the screen and get a feel for is there some thing happening in your mind when you're watching something on the screen where it feels like almost like mirror neurons you know I use that word loosely but mirror neurons are firing and you feel like there's a connection between you and that thing that step of it you definitely can I think imagine cuz you can experience that that's what we call open loop so what that means is that like if you imagine this like cartoon loop of like there's a human in there the human has some intent in their mind uh that intent goes down into some action and then they observe the outcome of that action and then that goes back into their brain. So it's like there's this loop control.
开环就是指回路是开放的。他们只是在观察,但看不到屏幕上的任何结果,纯粹是在观察发生了什么。闭环则是你能实际看到结果,然后这个结果会反馈回来,闭合回路。在开环阶段,你当然可以运用自己的经验和直觉来创造一种感觉适合你设计任务体验。但闭环极其困难,因为你最终无法直接体验神经接口。例如,我们有多种模拟这类控制的方法。当我移动光标时,我们有一个模型会读取移动速度,然后将其转换成某种脉冲表示,就像这种移动的神经表征可能的样子。但如果你真的用这个来优化界面的某些方面,它会失败,因为你移动的是一个真实的光标。所以你有所有这些额外的反馈机制:手在空间中的位置、移动物理鼠标时的摩擦感、触控板上的摩擦感、点击的压力感。你无法真正感受到使用它的感觉。而这实际上代表了那些没有这些东西、只能看着屏幕的人的感受。所以闭环方面肯定不是你能直接共情的。你可以勉强尝试,但它从来都不是可靠的感觉指标。但为了达到那个目标,我们知道他们必须第一次连接到植入物,而且之后还需要校准模型。所以我们需要为开环方面准备一些东西,并且需要基于学术界和其他研究工作做出最佳猜测,来推断闭环应该是什么感觉。我们还希望他们能在实验室之外、在我们不在场的任何会话之外实际使用它。真正的魔力都发生在外面,当我们不在房间里,他们只是用脑机接口做事情的时候。我们希望这能成为他们真正可以日常使用的东西。所有的工作都围绕着让这个东西可靠、拥有足够简单的界面,让他们能在 Mac 上实际使用,再次浏览、点击、操作电脑。所以我们很清楚,这不再是一个 iOS 应用了。它将只是一个 Mac 应用,让他们能重新使用电脑,我们也清楚了体验的基本构建块是什么。
Open loop just means the loop is open. They are observing but they don't actually see any outcome on the screen. They're just purely observing what's going on. Closed loop is when you can actually see an outcome and then that obviously feeds back and closes the loop. The open-loop stage, you can definitely use your own experience and your own instincts to create an experience that feels right for the task you're designing for. But closed loop is extremely hard because you ultimately cannot experience the neural interface directly. For example, we have various ways of simulating these kinds of control. When I move the cursor, we have a model that will read the velocity of that movement and then convert it into some spike representation, like what the neural representation could look like for this kind of movement. But if you're actually using that to optimize some aspect of the interface, it's going to fall flat because you're moving a real cursor. So you have all of these additional feedback mechanisms from where your hand is in space to the feeling of the friction when you move a physical mouse or the friction on a trackpad to the pressure of a click. You're not going to get any real sense of how that feels to use. That is actually representative of what it would feel like for somebody who doesn't have any of those things and is just looking at what's on screen. So the closed-loop side is definitely not something you can empathize with directly. You can weakly try, but it's never a reliable indicator of what that feels like. But to even get there, we knew that they would have to connect to an implant for the first time, and beyond that, we knew that they would have to calibrate a model. So we would need something for the open-loop side, and we would need to have some best guess based on academia and other research work as to what that closed-loop thing should feel like. And then we also knew that we wanted them to actually use this outside of a lab, outside of just any sessions with us. The real magic is all going to happen outside when we're not in the room and they're just using their BCI to do stuff. We wanted this to be something that they could actually live with. And all of the work that goes into making something reliable, with a simple enough interface that they can actually use on their Mac to go around, click on stuff, do stuff with their computer again. So we had a really good sense that it's not going to be an iOS thing anymore. It's going to be just a Mac app that enables them to use their computer again, and what the basic building blocks were going to be of that experience.
我想深入探讨其中一些构建块,也许第一个我们可以谈谈的就是光标。作为设计这种体验的人,你需要考虑哪些事情?因为我的假设是,有很多东西我们认为是理所当然的,而传统的 B2B SaaS 设计师开箱即用。
I want to drill into some of those building blocks and maybe the first one we could talk about is just the cursor. What are all of the things that you have to think through as someone that's designing this experience? Because my assumption is there's so many things that we take for granted and traditional B2B SaaS designers just get out of the box.
我认为光标最终是体验的焦点。当我们使用光标时,它是你在二维空间中意图的焦点。所以它涵盖了很多东西。B2B 的人免费获得的东西,实际上任何为光标构建界面的人都免费获得的东西,是使用光标时涉及到的极其丰富的感官体验,这些体验现在已经如此潜意识化,以至于你根本不会去想它。但这就是我之前提到的所有额外感官:摩擦、压力、声音。所有这些意味着拥有一个感觉像自己延伸的光标,你只专注于光标在做什么。你不再过多考虑手在做什么。你在构建应用时不需要考虑这些。你在屏幕上放一个按钮,可以给按钮添加一些深度,做那些让按钮点击感觉很好的事情。但最终,当你这样做时,它考虑到了所有这些其他感官媒介。所以我们的挑战是,我们的参与者不会有这些。我们仍然希望光标——因为它是体验的焦点——感觉很棒。那么,以什么正确的方式将某种形式的反馈带到光标上呢?有一百万种方法可以做到这一点。还有一百万种我们甚至还没有原型或尝试过的方法。但我们的第一个尝试实际上不是我们今天拥有的光标版本。我们的第一个尝试是某种常规光标,当你连接到植入物时,它会进行一个疯狂的过渡。它会接管你的 Mac OS 光标,然后旋转成这个新的神奇脑机接口光标。然后主要问题是,我们有一个模型(我们稍后会谈到),它会将一些意图转化为点击的概率。所以会有左键点击的概率和右键点击的概率。移动没问题。当你移动光标时,这是你能看到的。但移动和点击的组合呢?如何让这感觉不像一个离散事件(点击就会发生),而是一个连续的交互?你希望它感觉连续的原因是,在目前阶段,点击存在一些延迟。我们的模型需要一些时间来提升其置信度,认为参与者确实试图在这里点击。这意味着如果模型完美,每次延迟都会相同。实际上延迟就是植入物的采样频率,目前大约是 15 毫秒。所以每 15 毫秒,计算机会获得关于你神经活动的新信息。在完美世界中,点击会在你想到点击的那一刻发生,在 15 毫秒内,你实际上不需要 UI,因为模型总是正确,延迟基本为零。点击会在屏幕上发生。挑战在于存在一些延迟。假设是 100 毫秒、200 毫秒、300 毫秒,在这个范围内。它分配的概率有时会在 100 毫秒后达到正确概率,有时会是 300 毫秒。
I think the cursor is the focal point of that experience ultimately. When we use a cursor, it is the focal point of your intent in two dimensions. So it encapsulates a lot. The thing that B2B people get for free, and really anybody building an interface for a cursor gets for free, is that there's such a rich spectrum of sensory experience that just goes into using a cursor that at this point is so subconscious that you don't really think about it. But it's all of those additional senses I talked about earlier: friction, pressure, sound. What all of that means is having a cursor that feels like an extension of yourself, and you're just focused on what the cursor is doing. You don't think too much about what your hands are doing anymore. You don't have to think about any of these things when you're building an app. You put a button on screen, you can add some depth to that button, and do the really nice things to make a button click feel great. But ultimately, when you do that, it's taking into account all of these other sensory mediums. So our challenge was that our participants are not going to have that. We still want the cursor, because it is the focal point of the experience, to feel amazing. So then what's the right way to bring some form of that feedback to the cursor? There are a million ways to do this. There are a million ways that we still have not even prototyped or tried. But our first take on this is actually not the version of the cursor we have today. Our first take was some form of a regular cursor that you see when you connect to your implant, it does this wild transition. It'll take over your Mac OS cursor and spin into this new magical BCI cursor. And then the main question was, we have a model, we'll talk about that later, that will take some intents into a probability of a click. So there'll be some probability of a left click, some probability of a right click. The moving is fine. When you're moving a cursor around, that is something you can see. But it's the composition of moving and clicking. How do you make that feel not like a discrete event, like the click will just happen, but a continuous interaction? The reason you want that to feel continuous is because right now, at the current stage, there is some latency to that click. It takes some amount of time for our model to ramp up its confidence that the participant is trying to click here. And what that means is that if the model were perfect, it would be the same amount of latency every time. And actually the latency would just be whatever the sampling frequency of our implant is, which is like 15 milliseconds right now. So every 15 milliseconds, the computer will get some new information about your neural activity. In a perfect world, that click would happen the second you think about the click, within 15 milliseconds, and you wouldn't really need a UI here because the model is always right and the latency is basically zero. The click will happen on the screen. The challenge is that there is some latency. Let's say it's like 100 milliseconds, 200, 300, somewhere in that range. That probability it assigns is sometimes it'll be 100 milliseconds before you get to the right probability. Sometimes it'll be 300.
尤其是当我们思考 BCI 旅程的第一个月会是什么样子时。它很有可能偶尔会出错——比如在你没想点击的时候点击了。这显然是我们想要解决的问题,但我们希望至少能让这种体验可见,更重要的是可预测。所以当你第一次尝试点击时,它不是那种随机事件——过一会儿点击就会在屏幕上发生——而是一种连续的交互,就像你把手指按在触控板上一样。那里有一种连续的交互,你可以看到强度随时间逐渐增加。我们的直觉是,这在视觉上会比点击突然发生而没有任何反馈要好得多。所以第一个版本使用颜色来执行点击。我们用了非常漂亮的蓝色表示左键点击,橙色表示右键点击。实际上我们会根据每种点击的相对概率混合这两种颜色。我们还使用了深度和光标轻微的透视变化。所以当你点击时,光标会稍微向内倾斜。倾斜的程度由左键点击的概率驱动。蓝色的鲜艳度与填充程度相关。它会从光标的底部开始,然后随着倾斜逐渐向尖端延伸,从浅蓝色变成非常亮的蓝色。我们还可以利用显示器的 HDR 部分让它更亮一些。但你不希望这样,因为你每天要点击一千次,所以你需要它很微妙。我们不希望它感觉像是一个沉重的交互。但直觉是,让它连续并直接映射到解码器的输出,第一,能让用户感觉像流畅的交互而不是离散的事件;第二,当事情不顺利时,能给我们和用户一些可见性——你不希望屏幕上随机出现点击而你完全不知道原因;第三,一个更微妙的点是,神经接口的有趣之处在于——这对所有接口都成立——你也在学习如何使用它。所以当你真正处于闭环中时,流程的最后一步是你观察输出,然后根据输出改变你的输入。因为这是一个控制回路,看到管道的每个阶段——从 0.1 概率到 0.2 再到 0.3——能让你在潜意识中随着时间的推移学会调节自己的行为以更好地点击。如果模型完全错误,那对你来说会非常困难。但如果输出有细微的不准确,就会发生一种共同适应。所以我们希望确保运动也是如此。例如,你可以学会随着时间的推移更精确地移动它。这个学习步骤——想象一下如果你今天有了一条新手臂,最初的动作会非常奇怪和不协调。但随着时间的推移,你会学会精确地调节手臂的每一部分。但这确实是一个挑战,就像学骑自行车一样。我们希望运动中的那种连续性也适用于点击交互。
And especially when we're thinking about what the first month of the BCI journey would be like. There's also a very good chance that sometimes it'll just be wrong. So it'll click when you didn't intend to click. That's obviously something we want to solve for, but we want the experience of that to be at least visible and much more importantly to be predictable. So when you are in this flow of trying to click for the first time, it isn't this random thing that after some amount of time the click will happen on screen, but it's this continuous interaction much like when you press your finger down on a trackpad. There's a continuous interaction there and you can see the intensity of that ramp up over time. Our intuition was that that would just feel a lot better visually than just the click happening with no feedback at all. So the first version of this used color to perform that click. We had, I think, this really nice blue for a left click and this orange for a right click. And we would actually mix the two depending on the relative probability of each click. And then we also use depth and a slight perspective shift on the cursor. So as you were clicking, the cursor would sort of tilt inward just a little bit. The amount of tilt was driven by the probability of the left click. And the vibrancy of the blue was attached to that alongside like how much it filled in. It would start kind of at the base of the cursor and then as it was tilting in it would just sort of grow or stretch towards the tip and go from this like a nice light blue to a very bright blue. We can also use the HDR parts of the display to make that a little bit brighter. You don't want this because you're going to click a thousand times a day. So you do want this to be subtle. We didn't want this to feel like a weighty interaction. But the intuition was that having it be continuous and map directly to the output of the decoder is something that one, let them just feel like a smooth fluid interaction as opposed to just this discrete thing. Two, this is always a challenge: when things aren't working, give us some visibility, both them and us some visibility into that. You don't want just random clicks happening on screen and you have no idea. And then the third thing is a more subtle point, but the interesting thing about a neural interface — this is true about all interfaces — but you are also learning how to use it. So when you are actually in a closed loop way, again the last step of that flow is that you actually observe the output and then change your inputs based on the output. So because it's a control loop, seeing each stage of that pipeline when it goes from 0.1 probability to 0.2 to 0.3 gives you a way to subconsciously over time learn to modulate your own behavior to click better. So if the model is just completely wrong, this is going to be a very, very hard thing for you to do. But if there are subtle inaccuracies in that output, there's kind of a co-adaptation that can happen over time. So we wanted to make sure that that's true of motion. For example, you can learn to move it more precisely over time. And that learning step — obviously imagine if you had a new arm today, the very first things you would do would be super weird and uncoordinated. But over time you would learn to really precisely modulate every single piece of that arm. But that is a challenge, right? It's like learning how to ride a bike. We wanted the same sort of continuity that you have in motion to apply to the interaction of clicking.
你从参与者那里学到了哪些经验,从而改变了你对点击的看法?比如你现在处于什么阶段,造成这种差异的原因是什么?
What were some of the lessons that you were learning from participants that evolved the way you thought about a click? Like where are you at now and what is the reason for that delta?
最大的变化实际上是交互空间发生了很大变化。所以我们从第一天只想要一两个点击,发展到更丰富的计算机操作集合。滚动是一个很大的方面,它与移动光标非常不同,比如你实际如何滚动内容。当然还有拖拽——拖拽看起来应该很简单。拖拽是长时间按住的操作。当我们按住时,实际上会感受到一些阻力,比如开关内部弹簧的阻力,这给了我们继续施压的信号,而对于没有阻力的东西来说,这种信号并不完全存在。从表征上看,如果你观察神经数据,有时看起来更像是先过渡到按下状态,再过渡到抬起状态。这可能更适合作为拖拽光标的模型——就像你进入按下状态,然后进入抬起状态。但另一方面,如果走这条路,那么其他非拖拽的点击就会变慢很多,因为你必须在两种状态之间切换。也许有更好的建模方法,这确实是很多人花时间思考的问题。我们基本上希望第一天你就能移动和点击,既然能做到这一点,我们也希望至少再有一种点击。这只是一个很好的功能。坦白说,右键点击对于弹出菜单很有用,但不是必需的。更多是因为我们想测试在混合多种点击的情况下,你是否能可靠地在它们之间切换。但随着时间的推移,使用计算机需要更丰富的交互空间。你叠加得越多,那种简单的模型——你看到光标内混合两种点击——就越行不通。我们的第一位参与者,大约一个月后,信号质量与第一天大不相同。所以我们退了一步,也支持了停留交互,因为我们发现在这种模式下,点击实际上比运动更难解码。那种在移动光标中看到每一个细微动作的连续能力,让他移动光标比实际执行点击更容易。所以这意味着我们想要一个从移动光标开始的梯度控制,如果你能移动光标,就可以用运动作为点击的信号。这就是停留的作用。如果你放慢速度并停下来,就可以在某个位置保持,然后以这种方式执行点击。我们想支持这一点。
The biggest thing was actually that the space of interactions changed a lot. So we went from let's say just one click or two clicks that we wanted for the first day to the richer set of all of the things you do on a computer. So scrolling is a huge thing that's very distinct from how you move a cursor, like how you actually scroll something. Of course, dragging — dragging is something you think should just work. A drag is something that you hold down over time. When we hold it, there's some resistance that we actually feel, right, from the literal spring inside of these switches that is kind of our signal to keep applying pressure that isn't fully there with obviously something that offers no resistance. And it's also representationally: if you actually look at the neural data, what you see is sometimes what it looks like is actually more of a transition into clicking down and then a transition back into clicking up. That may be a better model for a cursor that drags — it's almost like you get into the click down state and then you get into click up. The flip side is that if you go down that route, every other click that isn't a drag gets a lot slower because you now have to transition between these two states. There may be a much better modeling approach here, and that's definitely something that a lot of people spend a lot of time thinking about. We wanted basically on the first day for you to be able to move and click, and because you could do that, we also wanted to at least have one other click in there. It's just a nice thing to have. Frankly, right click is a pretty useful thing for just popping over something, but not necessary. It was more just because we wanted to test out if you have multiple clicks in this mix, can you reliably switch between them? But over time, then there's a far richer space of interactions that you need to just use your computer. And the more you stack in there, the less having just one simple model where it's just like you can see what the thing is and it's mixing between the two inside the cursor — that just kind of fell apart. Our first participant, about one month into his journey, we had a very different signal quality than we did on the very first day. And so we actually took a step back and also supported a dwell interaction because we found that the clicks in this regime were actually much harder to decode than the motion. Something about that continuous ability to see every slight movement you make in a moving cursor made it easier for him to move the cursor than it did for him to actually perform a click. And so what that meant is we wanted this gradient of control that goes everywhere from just moving the cursor, and then if you can move the cursor, you can use motion as your signal for when to click. So that's what the dwell would do. So if you slow it down and bring it to a stop, then you can basically hold that in a certain place and then perform a click that way. We wanted to support that.
然后我们还希望支持左键点击、右键点击、拖拽、滚动、缩放以及随着你逐步提升所需的所有其他交互。我们从最初一个非常简单的、仅用颜色显示两种点击的光标,演变为一种更圆形的表示形式。这就像一个圆圈加一个点,几乎像瞄准线。这实际上让你更容易看出点击概率的差异。在这个世界里,我们不再使用颜色,而是只用了圆的外半径。当你瞄准某个东西、意图点击时,它会缩小并聚焦到那个物体上。基本上,随着你点击概率的提升,它会收缩成一个点。也许因为我们的眼睛对运动比颜色更敏感,这反而更容易用来执行点击。而且它也更能自然地支持悬停。圆形光标实际上比现在的普通光标更容易看穿。而且由于你主要把它当作运动信号的指示器,我们可以把它做得稍大一些,而不会遮挡你实际想点击的内容。所以我们从这种更传统的指针样式转向了这种圆形样式。这也为在不同交互模式之间更轻松地切换打开了大门。
And then we also wanted to support layering in a left click, a right click, a drag, a scroll, a zoom, and all of the other interactions that you need as you progressively ramp up. We went from just having this really simple cursor that could show you just using color, these two clicks, to a more circular representation. So this is like a circle and a dot, like a reticle almost. This actually made it a little bit easier to see the difference in your flick probability. So in this world, we didn't use color anymore, but we used just the outer radius of the circle. As you were targeting something, with the intent to click, it would scale down and focus in onto that thing. It would collapse to a point basically the more you ramp up your click probability. Perhaps because our eyes are more sensitive to motion than color, this was an easier thing to actually use to perform a click. And it also supported just using dwell more naturally. The circular cursor was easier to actually see through than a normal cursor is today. And because primarily you're just looking at it as your signal of where the motion is, we could make it a little bit bigger without it obscuring stuff you actually want to click on. So we went from this more traditional pointer style to this sort of circular style. And that opened also the door to going between different interaction modes more easily.
所以我们最终构建了一个不同的模式切换器:你可以把光标猛地甩到屏幕右侧,然后它会进入轨道,你用移动光标的同一个模型来选择新模式,再退出来。也就是说,你向右甩、上下滚动、向左甩。它的好处是,如果你学得好,操作可以非常快,不像普通悬停那样需要花时间选择模式——它只是一个流畅的动作:向右移动到屏幕的某个区域(通过神经肌肉记忆你学会那里对应什么模式),然后弹回左边。一旦我们有了不同的模式,比如拖拽、滚动,我们就可以轻松改变这个瞄准线的行为及其视觉外观,让你清楚知道当前模式下它会做什么。我认为,很多这类操作如果用传统光标样式、把所有东西都塞在光标里,会更难实现。这里的权衡是——随着模型改进,我们肯定会回头重新审视——你确实需要切换模式来改变交互方式。
So, we ended up building a different mode switcher where you could sort of slam your cursor to the right side of the screen and then it would go on rails and you would use the same model that you used to move the cursor to pick a new mode and then come out. So, you would shoot right, scroll up or down, shoot left. What's nice about this is that it was something that could be really fast basically if you learn to use it well as opposed to something where you have to use your normal dwell because then there's no time spent actually selecting that mode and it's just one fluid motion of going to the right to the part of the screen that you learn through this sort of neural muscle memory of where this mode should be and then popping out back left. And then once we had different modes, we had like drag in there, we had scroll in there, we could easily change the behavior of what this reticle is going to do and the visual appearance of what it looks like to make it obvious to you what it's going to do now that you're in this mode. A lot of that was harder to do, I'd say, with just the conventional cursor style and having everything exist in one space within the cursor. The trade-off there, and this is definitely as our models improve, something that I think we'll then go back and revisit is you do have to switch modes to actually switch the way in which you do your interaction.
所以,BCI 的世界很有趣:如果模型完美,它总能做你想做的事。这听起来是句废话,但在那个世界里,你根本不需要这些界面。但我们赌的是,到达那个世界的方式更像是沿着梯子一步步走,而不是试图一开始就搞定一切。我们基本上希望你现在就能用光标做所有这些事情。即使我们实现的方式更多是设计和工程优化,而不是我们认为的完美最终方案。
So yeah, the world of BCI is like an interesting one where if the model is perfect, it's always going to do what you want. It's like an obvious statement, but in that world, you really don't need these interfaces. But our bet is that the way in which you get to that world is more of walking along this ladder than it is like trying to do everything at once up front. We want basically you to be able to use your cursor to do all this stuff today. Even if the ways in which we get there are more of design and engineering optimizations than they are what we think the perfect final solution is.
你预想的梯子上有没有不同的阶梯?比如你特别期待为这个产品带来的具体障碍或能力解锁?
Are there rungs on that ladder that you anticipate said differently? Maybe like specific barriers or unlocks in the capabilities that you're really excited to bring to this product.
如果可能的话,梯子上我最兴奋的一级——这绝对是个开放问题——就是我们能不能删掉光标?想想光标归根结底是什么:它只是你意图的焦点,你必须移动它。你做的 90% 的事情更像是“我想和屏幕上的这个东西交互”。你需要光标吗?如果你能读懂我想和屏幕上这个点交互的意图,可能不需要。我认为光标在直接操作任务中非常棒。当你真正想直接操作某个东西时,光标会淡出,感觉很好。你会觉得“我在旋转这个东西”、“我在放大它”。无论施加什么变换,光标都不存在了,只有我的手在和屏幕上的数字元素交互。这是这个新世界中的一个关键交互模型。但第一步——90% 的时间我只是想和某个东西交互——我认为不需要光标。所以那是梯子上的一级,而且还有好几级要走。
The rung on the ladder that I'm most excited about, if this is something that we can do, I think it's definitely an open question, would be can we delete the cursor? So, if you think about what the cursor is ultimately, it is just this focal point of your intent that you have to move around. 90% of the stuff you do is more just like I want to interact with this thing at this point on my screen. Do you need a cursor for that? If you can read that intent that I have that I want to interact with this point on my screen, probably not. I would say the cursor is amazing for direct manipulation tasks. So when you actually want to directly manipulate something, the cursor kind of fades away when you do that, which is really nice. And you can just feel like I'm rotating this thing. It feels like I'm scaling it up. Whatever transformation I'm applying to it. It feels like the cursor isn't there anymore. It's just my hand interacting with this digital element on screen. That is a key interaction model in this new world. But like if that first step of just the 90% of the time I'm just trying to interact with something, I don't think I need a cursor for that. So that's one step of this ladder I would say that is quite a few rungs out.
在通往那个目标的路上,我个人非常期待一个不需要切换模式的世界。某种我们最初的设计形式:你不仅能看到光标内的点击概率,还能看到你想如何进行交互。举个老套的例子:如果我想放大某个东西,视觉上我们仍然想给你一些关于缩放的连续反馈——光标几乎像减数分裂、细胞分裂一样,可以爆裂成两个手指点,然后这些点就成为你进行缩放交互的视觉锚点。那是梯子的另一级。如何超越当前的交互模式,进入一种不需要切换模式的新光标样式?这种模式切换其实可以被我们的模型学习。我们捕捉你的意图,不是那种“我想和屏幕上这个东西交互”的高阶意图,而是具体的“我想拖拽”。那可能会感觉很神奇,因为当你想拖拽时,它就直接切换到拖拽模式。你不需要通过我们当前的方式(模式切换器,或者用一次点击快速切换到特定应用的首选模式)来传达这个意图。这当然在功能上已经达到了 90%,因为大多数时候你的第二个交互就是拖拽。我很好奇你的会是什么,但我的肯定也是这个,因为你可以按应用来设置。
On the way to that, I'm personally very excited about a world in which you don't need to switch modes. Some form of that first design that we had where you can see not just the click probabilities inside of this cursor, but the way in which you want to do this interaction. So for example, if it is something like I'll give you one cheesy example like if I want to zoom in on something right visually we still want to give you some form of continuous feedback about that zoom the cursor almost like meiosis like cellular division could explode into like two finger points right and then those could then be your visual anchor for how you're doing this zoom in interaction. That's another step of a ladder. How do you go beyond basically the current interaction model into something that doesn't require switching modes to a new style of cursor? That mode switching is something that our model can actually learn. So we pick up on your intent, not in this case of just this crazy high order intent of I want to interact with this thing on my screen, but rather the specific I want to drag. Now that could feel like magic because it's just like when you want to drag, it just switches to the drag thing. You don't have to channel that intent through how we currently do it, which is we have a mode switcher and we also have a quick switch where you can use one of your clicks to quickly switch to your preferred mode in this specific app. So that gets us of course functionally 90% of the way there because most of the time your second interaction is dragging. I'm curious like what yours would be, but that's definitely what mine would also be because you can do it on a per app basis.
就像在 Illustrator 里你大部分时间都在缩放,对吧?你可以轻松地重新映射,功能上能覆盖 90%,但它仍然缺少那种神经指向设备独有的魔法,这些物理指向设备做不到的,那就是我们能读取某种底层意图。我认为接下来的 10 个台阶就是要把真正的原始意图带入我们的体验。所以到目前为止的目标就是让我们能够以同样的保真度使用光标。我认为第一位参与者创下的纪录设定了很高的门槛,大概是 9.5 bps。BPS 是衡量我们能读取的信息量的指标。你可以把它看作一个分数,因为你可以用同样的任务来测量你光标的 BPS。比如用触控板或鼠标,点击屏幕上的位置,它测量你完成的速度和准确度。第一位参与者创下了这个纪录,因为他非常喜欢使用神经光标。他每天使用,持续了几个月,变得越来越熟练,从大约 4 bps 开始,最低是 2 bps,然后一路提升到 9.5 bps。作为参考,我大概能做到 10 bps,那是我个人的纪录。我可能用的是触控板。但主要目标是让光标的使用达到与现有设备相同的保真度。这让我们如此兴奋的原因是,光标不仅是你操作的核心,也是你使用电脑的方式。想想这意味着什么:我们用电脑来提高效率、表达自己、与他人交流、以及娱乐。一个流畅的光标就能开启人类体验中相当大的一部分。所以这绝对是我们的目标,也是为什么我们在不同的权衡中探索,只是为了实现某种形式的光标使用。因为我们知道还有更好的东西在等着我们,但之所以不重新发明轮子,是因为轮子是你上高速公路的方式,也就是现有的系统。这绝对是重点。但在此之上还有更高的台阶,那些不仅仅是投射到我们今天使用的光标上的东西,而是 DCI 独有的。
Like if in Illustrator you mostly are zooming, right? You can easily remap that and functionally that gets you 90% of the way there, but it's still lacking the actual magic that is unique to something like a neural pointing device that these physical pointing devices just can't do, which is that we can read some form of our underlying intent. I think the next 10 rungs on this ladder are just bringing the actual raw intent into our experience. So the goal so far has been just to enable the use of the cursor with the same fidelity that we can use it. I think the current record our first participant set a wild bar. It was like 9 and a half bps. BPS is just a metric for the information that we can actually read. Really you can just think of it as a score, because you can use the same task that they use to measure the BPS of your cursor. Like you can use it with a trackpad or a mouse and just do what they do, which is click on positions of the screen, and it measures basically how quickly you can do that and how accurately. So yeah, first participant set this wild bar because he just loved using his neural cursor. He used it every day for months and got insanely better at using it, and went from something like 4 bps, at the very bottom rung was like 2 bps, and he just walked his way up this ladder to I think 9 and a half is where he ended up. For context, I can do about like 10. That's my personal record. I'm probably on a trackpad. But the primary goal is to enable the use of the cursor at the same fidelity that we can use it. The reason this is so exciting for us is that the cursor, beyond just being the focal point of your agency, is also how you use a computer. So if you think about what that entails, we use computers to be productive, to express ourselves, to communicate with each other, and to have fun. It's a pretty wide subset of the human experience that can be enabled with just a really fluid cursor. So that is definitely the goal, and that's why we walk this space of different trade-offs just to enable some form of cursor use. Because we think there's obviously this better thing out there that is in the back of our minds of where we can go from here, but the reason not to reinvent the wheel in this case is just because the wheel is how you get on the highway, which is this existing system. That's definitely the focus. But there are rungs beyond that, which are things that are not just a projection onto a cursor that we use today, that are very unique to DCI.
我想了解更多关于产品体验本身是如何随着时间演变的。随着参与者使用越来越多,他们的行为如何影响了一些实验,以及你们是如何思考产品应该成为什么样的?
I'm interested in learning more about just how the product experience itself has evolved over time. So as you're getting more and more usage from these participants, how are their behaviors shaping some of the experiments and just how you were even thinking about what the product needed to be?
我们在系统中内置的一个备用方案是语音输入。只需说一个简单的命令就能重新校准模型、更改一些参数——这不是我们设想的最终体验,但在早期事情远不那么可预测时,它是一个非常可靠的备用选项。我们的第三位参与者实际上患有 ALS,这意味着他无法说话,所以语音输入在这里行不通。这实际上渗透到了我们界面的每一个角落。那些我们习以为常的语音输入,我们不得不更仔细地思考一个好的备用方案是什么。但我觉得最酷的事情直接来自第三位参与者,就是我们称之为“停车位”的功能。光标总是会有些移动。随着模型改进,希望这不会发生。但短期内,你仍然想使用它。所以会有一些移动是模型错误猜测的,或者与用户实际思考相关,比如当他们和旁边的人说话或看 YouTube 视频时。模型会错误地将这些活动标记为移动。你可以在建模方面解决这个问题——长期努力——但短期内,你需要一种方法来停靠光标。之前的参与者可以说“关闭光标”,它就会关闭,然后说“打开”就能重新开启。但对于第三位参与者,显然他做不到。所以我们需要想办法用光标来关闭光标,再用光标来重新开启。我们尝试了几种方法,但最终确定的是“停车位”功能:他们可以把光标“吃掉”或射到屏幕右下角,如果推到那里,它就会停靠。一个小东西会弹出来,锁定光标,然后他们可以在那个表面内使用手势把它拉出来。我们应用了一些小变换来保持它静止,并在里面模拟了重力。他们仍然可以通过用力推或通过特定的模式(比如点、点、点)自己把光标完全拉出来。这非常酷,因为我们主要是为第三位参与者做的。但当我们发布这个功能时,前两位参与者说:“等等,这太棒了。”然后他们在最初几周使用得比第三位参与者还多。从每个参与者身上都能学到很多,而且大多数时候,他们反馈的东西往往也会惠及其他所有人的体验。
One fallback we built into the system in general was voice input. So just being able to say a simple command to recalibrate the model, to change some parameters of the model — not something we envision as our final sort of experience, but it was a great in those early days when things were far less predictable, it was a really reliable fallback option for them to have. Our third participant actually had ALS, which means that they cannot speak, and that means that of course voice input is not something that's going to work well here. So this actually bled through to every edge of our interface in some sense. Things that we had taken voice input for granted for, we had to think a little bit harder about what a good fallback option would be. But I think the coolest thing that actually came out of this directly from our third participant was a feature that we called the parking spot. The cursor is always moving to some extent. This is one of those things in that bucket of as models improve, this will never be the case hopefully. But in the short term, you still want to use it. So some amount of motion is going to be there that the model's just kind of incorrectly guessing or is correlated to them actually just thinking when, for example, they're talking to somebody next to them or they're watching a movie on YouTube. There's just some amount of activity that our models are incorrectly labeling as motion here. You can solve this on the modeling side — longer term effort — but shorter term, it's like you just need some way to park the cursor. Our previous participants could just say "hey turn the cursor off" and then it would just turn off, and then they could just say "turn it on" to bring it back on again. For our third participant, obviously, they couldn't do that. So we needed to think of some sort of way to use the cursor to turn the cursor off and then use the cursor to bring it back on. We tried a few things here, but the final thing we landed on was this thing called the parking spot, where they could just sort of eat their cursor or shoot their cursor into the bottom right of the screen, and if they push it there, it'll park it. It'll like this little thing will pop out, it'll lock up the cursor, and then they can actually use a gesture within that surface to bring it back out. So we apply a bunch of little transformations in there to hold it still. We simulate gravity inside of it. They can still, by pushing really really hard or by actually going through a specific pattern, like a dot dot dot, they can pull the cursor back out fully on their own. So that was a really cool one because we did that primarily for the third participant. But when we shipped that, our first two at the time were like, "Wait, this is amazing." And then they actually were bigger users for the first few weeks — they were using it far more than our third participant was. There's a lot you learn from each individual participant that comes through, and most of the time actually what they give you feedback about often always bleeds out into everybody else's experience as well.
在光标上模拟重力是什么意思?
What does it mean to simulate gravity on the cursor like that?
你可以理解为我们在做正确的数学计算,但你可以想象当光标进入这个停车位时,它就像掉进了一个山坡。光标会滚下来,然后停在那里。我们可以调整山坡的深度,比如是一个陡峭的山谷还是一个平坦狭窄的小山。但实际上,这意味着他们需要更用力地推才能把它滚上去,并在上升过程中保持动量。模型在任何给定点的实际输出不是位置,而是速度。
You could think of it as we just do the right math to work this out, but you can think of it as when the cursor goes inside this parking spot, it's almost like it falls into a hill. So the cursor kind of rolls down here and then it's sitting there. We can obviously tune the depth of that hill. So we can tune if it's a really steep valley or if it's just a really flat narrow hill. But what it means in practice is that they have to push harder to roll it up there and maintain momentum as they go up. The actual output of the model at any given point is not a position. It's actually a velocity.
所以这其实就像是一种轻推。在这个场景下,意思是如果我们加入重力把光标往下拖,他们就得手动把它移出来。我说我们的第三位参与者在头两周其实用得不多,就是因为那个基于重力的版本对他并不奏效,这真是个有趣的深坑。他和很多晚期 ALS 患者一样,用眼动仪作为与世界沟通和使用电脑的主要方式。这意味着,比如他看电影时,眼睛会满屏幕乱扫。有趣的是,当他看电影时,即使有很强的重力,光标也会在他根本没看它的时候弹出来。但当他真的把目光聚焦在光标上,试图在强重力下把它移进停车位时,又很难把它推上去。所以非常有意思。
So, it's actually like a nudge, so to speak. What that means in this context is like if we add gravity to drag the cursor down there, they'll have to actually roll it out kind of manually. When I say that our third participant actually wasn't using it as much the first two weeks, it's because that version of it, the gravity based approach, didn't actually work for him, which is a really interesting rabbit hole. He, like a lot of people with late-stage ALS, used an eye tracker as his primary way of communicating with the world and using a computer. What that meant is that when he was watching a movie, for example, your eyes are all over the screen. It was a really interesting situation where when he was watching a movie, the cursor would just pop out when he's not even looking at it, no matter even with really really strong gravity. But when he really focused his eye on the cursor and tried to move it inside the parking spot with that high gravity, it was really hard to push it up. So really interesting.
是啊,这种需要深入思考的细节真的很有意思。
Yeah, that's a level of depth that you have to think through that is really interesting.
最终的解决方案是支持两种模式,可以用重力。对于另外两位参与者,重力模式效果太好了,以至于我们发现了其他问题——比如一位参与者说,他把光标停好后跟人聊天,结果屏幕就休眠了。他聊了整整 30 分钟,光标一动不动,而 Mac OS 如果光标不动,系统就会认为你没在操作,然后调暗屏幕并关闭。所以对他们来说,重力模式效果太好了。而对于第三位参与者,重力模式完全没按预期工作——当他看着光标时移不动,当他看向别处时,比如看电影,光标速度又非常快。所以我们最终为他改用了一种手势方案,效果非常好。他只需要做上、下、左、右的动作,这些小点会依次亮起,然后他就能这样把光标拉出来。
The solution there actually ended up being it supports like two modes where you can use it with gravity. And that actually for our other two participants it worked so well that we found other issues where like one participant was like hey I parked the cursor and I was talking to somebody and like my screen went to sleep. So like he was having like a whole 30 minute conversation where the cursor doesn't move and on Mac OS if the cursor doesn't move your system's like oh you're not doing anything and it'll dim the display and then shut it off. So in their case it worked too well. In his case it like yeah it literally did not do with gravity at least what you'd expect at all where when he looked at it he couldn't move it and when he looked away he could have huge velocities as he's like watching something. So we ended up going with a gesture for him and that worked super well. So he could just do like an up, down, left, right. Um, these little dots would light up. It would follow them in order, and then he would just like pull the cursor out that way.
听你说话很有意思,因为你痴迷于这些交互中最微小的细节。我是说,从各个角度观察和思考光标,想尽各种办法让它可行。而最终目标基本上就是删除它——在某种程度上删除你做的所有工作,这是一种很少有人能体会的有趣张力。
It's fascinating to listen to you talk because you're obsessing over the smallest pieces of these interactions. I mean, observing and thinking about a cursor from every possible angle, all the different ways that we could make this possible. And the end goal is to basically delete it. to delete all of the work that you've done to an extent, you know, like that's a really interesting tension that not many people get to operate in.
我认为最终目标就是打造一个令人难以置信的体验,并且每一步都朝着这个方向努力。所以,如果我们想要达到一个可以删除光标的世界——不管那意味着什么——我们的直觉是,这需要大量数据,需要对这些机制在那个世界中如何运作有更深入的理解。想想都觉得有点科幻了。但衡量标准不再是所谓的“正常人”在电脑上能做什么。你已经把人与电脑交互的可能性提升到了很高的水平,对吧?想想我在触控板上操作的速度,在那个世界里会显得很原始。
I think ultimately the goal is to just build an incredible experience and to do that in like every step of the way. So like if the way in which we get to a world where we could let's say like delete the cursor like whatever that means. Our hunch is that getting there is going to require like a lot of data. It's going to require a much more robust understanding of how these mechanics work in that world. Starts to get very sci-fi to even think about. But the measuring stick is no longer what a quote unquote normal person can do on a computer. You know, you've blown the roof off of what is possible in terms of interaction with computers at a very high level, right? like thinking about how quickly I'm able to do something on a trackpad will feel archaic in that world.
目前,交互模型和我们使用的仍然非常相似,但没有理由必须如此。所以我认为在这个世界里会出现新的交互模型,它们会非常不同,不再需要我们用物理手去表达什么。我不知道会不会显得原始,但肯定会非常不同。希望它会更加自然,更加直接。作为可能比普通听众花更多时间思考这个特定世界走向的人,你有什么特别的想法吗?
Right now, the interaction model is again very similar to the ones that we use, but there's no like reason it has to be. So, I think there will be new interaction models in this world that are yeah, just very different than having to use our like physical hands to like articulate something. I don't know if it would be like archaic, but I think it would just be very different. Like it'll just be hopefully a lot more natural and hopefully a lot less indirect. anything specific that you find yourself thinking about as someone that probably spends more time pondering where this specific world is heading than the typical person listening for instance.
我确实有一些想法。需要说明的是,这绝对只是我个人感到兴奋的东西,不一定是我们正在做的。而且有趣的是,在 Nerling,你问不同的人会得到完全不同的答案。但对我来说,很多想法都追溯到 15 年前那项最初激发我对这个领域兴趣的研究——我认为我脑海中最丰富的东西是视觉意象。同时,无论是作为软件设计师还是年轻时想成为电影制作人,我认为电脑的魔力以及电脑的繁琐之处,都在于把脑海中的图像花很长时间——无论是在 After Effects 还是 Final Cut 中——转化为视觉作品的过程,而人们更容易理解图像,一图胜千言。所以展望遥远的未来,我最兴奋的是这一方面。这意味着什么?我们经常谈论弥合人与人之间的共情鸿沟,但我觉得这是一个更具体的版本:我可以向你展示我的感受,而不是告诉你我的感受。此外,对于人们想要创造的所有东西,我认为每个人都有且共享一个丰富的内在空间。我总觉得有趣的是,当有人对我说“我现在不太有创意”时,我会问:你晚上做梦吗?如果做梦,看看你的大脑能创造和生成什么。我认为未来这将是一个巨大的赋能工具:你可以坐在电脑前,在这个世界里,电脑上的模型可以代表你的意图——比如说,在遥远的未来,在一帧之内。我们今天所说的“氛围编程”,在未来,模型生成你所描述的产物所需的时间,我认为会缩短到一帧。而另一面是,瓶颈变成了你能否清晰地表达你想要的东西——在艺术和工作的所有领域。这将是一种非常激动人心且狂野的体验:你坐在电脑前,基本上和它一起做白日梦,屏幕上会弹出直接延伸你脑海想法的东西。如果这一切浓缩到一帧之内,会是什么样子?我真的不知道,但我认为那将是一个非常令人兴奋的世界。
I think there are a few for me. The caveat being that like this is definitely just like the stuff that I'm excited about and not necessarily stuff we're working on. And I also think like the fun thing is like at Nerling you'll get a million different answers based on who you talk to. But I think for me it really a lot of it goes back to that study from 15 years ago that first sparked my interest in this field which is like I think the richest things in my mind are visual imagery. I also think that like both as a software designer and as like a wannabe filmmaker when I was younger like that's where so much of the work so much of the magic of a computer is and also so much of the tedium of a computer is is in the this process of taking an image in your mind and spending a very long time whether it's in software whether it's in After Effects or Final Cut to articulate that into this visual thing that it feels like people just understand more easily a picture is worth a thousand words basically that side of things is I think what I'm most excited about, you know, looking out into the very distant future. What that would mean, we talk a lot about, you know, bridging this empathy gap between folks, but like that feels like a much more tangible version of that where I can show you how I'm feeling instead of telling you how I'm feeling. And then also just for all of the stuff that people want to make in the world. There's this rich interior space that I think everybody has and everybody shares. I always think it's like funny when some people tell me like, oh, like now I'm just like not very creative. Do you dream at night? And like if so like that like look what your mind is capable of making and generating. I think that is a hugely enabling thing in the future where you can sit down in a computer basically and in this world where a model on that computer can act on behalf of your intent in let's say you know far down the line a frame. So what we today call vibe coding you know in the future like the time that it takes for a model to actually produce an artifact that you described will I think come down to a frame and the flip side of that is then what is the bottleneck here it's you actually articulate what it is that you want in all domains in the arts and in work this will be a very exciting and wild experience where you can just sit down on a computer and basically daydream with it and things will pop up on screen that are a direct sort of extension of what you have in your head and what would that look like if that came down to just being one frame. I truly don't know but I think that that's a really really uh exciting world to be in.
我们正在开发一个名为“盲视”的项目,这是这个阶梯上的第一级。它旨在让失去视力的人重新看到世界。我们的植入体位于大脑的视觉区域,而不是像当前植入体那样位于运动区域。我们可以从他们佩戴的眼镜中读取他们正在看到的内容,然后在大脑的相应区域创建正确的刺激模式,以在脑海中重建某种形式的图像。显然,这非常了不起,但这也是一个极其狂野的设计空间。最初的版本就像雅达利游戏机相比 PlayStation 5。我们能够创造的视觉保真度,就电极数量而言,相对于视觉信息的丰富程度来说非常有限。所以这本身就是一个完全独特的设计空间:如何在这个领域中重建一个“逼真”的图像?
We are working on something called Blindsight, which is the earliest rung in this ladder. It's for people who no longer have vision, a way for them to see again in the world. Our implant sits in the part of the brain for vision, not for motor movement like our current implant. We can read from a pair of glasses they wear what they're seeing, and then create the right stimulation pattern in that part of the brain to recreate some form of that image in their mind. Obviously, this is incredible, but it's also a truly wild design space. The first versions will be like Atari compared to a PlayStation 5. The visual fidelity we can create, in terms of the number of electrodes, is very small relative to how rich visual information actually is. So that's its own completely distinct design space: how do you recreate an image that is true to life in this domain?
图像的哪些特征是需要在这里强调的?你基本上想给人们哪些控制旋钮?因为他们应该对如何看世界有一定的控制权。这可能像科幻小说一样酷,比如能够放大物体,但也可能远比这更微妙,比如抖动。如果你熟悉抖动,那是 80 年代在低像素时代的一种方法,通过利用我们感知的伪影的算法,在较低分辨率空间中重建特征,仍然创造出阴影和纹理的感觉。那么,抖动在这个领域会是什么样子?
What features of that image are the right ones to highlight here? What knobs do you want to give people, basically, because they should have some form of control over how they see the world? That could be something as sci-fi and cool as being able to zoom into things, but it could also be far more nuanced, like dithering. If you're familiar with dithering, it was a way in the 80s, when we had low pixel counts, to recreate features through algorithms that exploit artifacts of our perception, using a lower resolution space to still create a sense of shadows and texture. So yeah, what would dithering look like in this domain?
我在这个节目上听过很多,但这是我想过的最有趣的设计机会空间之一。在你分享之前,它根本不存在于我的脑海中。这既惊人又引人入胜,不仅因为它的新颖性,还因为你们正在产生的影响。我知道我在开始录音前提到过,但我正在阅读一些参与者的故事。第一位参与者 Nolan 正在重返学校,并在网上找到了一份工作。这深深打动了我。你们所做的一切简直是最鼓舞人心的。
I've heard a lot on this show, but that's one of the most interesting design opportunity spaces I've ever thought of. It just didn't exist in my brain until you shared that. It's amazing and compelling, not just because of its novelty but also the impact you are having. I know I mentioned this before we started recording, but I was reading some of the participant stories. Nolan, one of the first participants, is going back to school and got a job on the internet. It moved me deeply. What you all are doing is about as inspiring as it gets.
我完全同意。我认为这正是我们团队乃至整个公司的人们为之奋斗的动力。我们非常荣幸能够身处这样一个世界,现在有 10 多个人正在实际使用这个设备。我们称他们为“神经宇航员”,就像宇航员一样。这是一个我们共同探索的全新空间。他们做了很多工作,比如回到光标最初的样子以及系统如何运作。现在情况发生了巨大变化,因为他们能够告诉我们感觉如何,哪些部分需要优化,哪些部分并不重要。最终,这对他们生活的影响对我们所有人来说都是极大的鼓舞。
I fully agree. I think this is the stuff that people get out of bed for on our team and throughout the company. We are so privileged to be in this world now where we have 10 plus folks who are actually using the thing. We call them the neural knots, like astronauts. It's a completely new space we're exploring together. They've done a lot, like back to that point of what the cursor used to look like and how things work. Now things have changed dramatically once they could tell us how things feel and what pieces we need to optimize, what pieces don't really matter. Ultimately, the impact it can make on their lives has been super inspiring for all of us.
假设有人在听,并且受到这段旅程的启发,他们想参与其中。你们已经开放了第二个设计工程师类型的职位。那么你能多分享一些这个职位会是什么样子吗?我肯定有人在听,他们会想:“是的,那可能真的很酷,但我也不是 100% 清楚那会是什么样子。”
Let's say that somebody is listening and they're inspired by this journey, they want to be a part of it. You all have opened up a second design engineer type role. So could you share a little more about what that role will be like? I'm sure someone is listening and they're like, 'Yeah, that could be really cool, but also I don't 100% get what that would look like.'
这个职位会涉及很多我今天谈到的工作,而且会更多。我们现在处于这样一个阶段:有 10 多个人每天都在积极使用这个设备。他们每周总共使用数百小时。一部分工作是继续为他们提供出色的体验。因为我们团队非常小,我们非常专注于把体验的哪些部分做到最好,而且我们对“好”的标准非常高。所以一部分工作就是让这个人们每天使用的东西变得出色,并持续保持出色。另一部分工作是阶梯上我们尚未探索的所有层级。现在我们有能够实际尝试并反馈的人,这个过程需要你从头到尾地完成。这意味着你可能对事情如何运作有一个初步的想法或直觉,即一个假设。从那个假设到让它成为别人能实际使用的东西,中间有很多实际工作。需要大量的设计工作和工程工作,把它变成他们可以首次使用和尝试的东西。微妙之处在于,如果某样东西需要他们积极使用和尝试,那么他们的第一次体验不一定是最有信息量的。我们参与者的伟大之处在于,他们非常愿意接受某样东西并尝试使用。他们在使用一天后和使用一个月后给出的反馈截然不同。有时,仅仅为了获得有意义的信号来判断这是否是我们想走的方向,需要的不仅仅是一个原型。你需要超越那种在一次会话中勉强能用的东西,把它变成他们可以连续使用一个月的东西。现实是,绝大多数这些东西都不会成功。所以现在我们处于这样一个阶段:既有一组我们知道要做得好的核心功能,因为我们希望这个产品存在于世界上;也有很多关于它可能成为什么的蓝天空想。这需要具备在设计和工程上从头到尾完成的能力,同时也要对实验性质感到舒适,这意味着关于这个东西是否最终会发布会有很多不确定性。因为你希望有人能长期使用它,所以还需要投入大量工作,以某种形式真正发布它,让他们能够尝试。
The role would be a lot of what I've talked about today, just a lot more of that. We're at this phase now where we have 10 plus people actively using the thing every day. They use it for hundreds of hours a week collectively. One piece is continuing to make that a great experience for them. Because we are such a small team, we are super focused on what pieces of that experience we spend the most time making great, and we have a very high bar for what great is. So one piece is literally making this thing that people use every day amazing and continuing to make it amazing. A second piece is all the rungs in the ladder we have not explored yet. The process of doing that, now that we have folks who can actually try stuff and tell us, is really something you have to go end to end on. That means there may be an initial idea or hunch you have about how things might work, a hypothesis. There's a lot of actual work in between that and making it something somebody can actually use. There's a lot of design work and engineering work to turn it into something they can use and try out for the first time. The nuance is that first experience they have with it is not necessarily the most informative if it's something they have to actively use and try. What's awesome about our participants is they are so down to just take something and run with it. They will give you great feedback after using something for a day versus a month. Sometimes it's not just a prototype to get a meaningful signal on whether this is a direction we want to go in. You need to go beyond something that will work in a hacky way in one session and turn it into something they can live with for a month. The vast majority of those things will not pan out, that's the reality. So right now we're in a phase where there's both a core set of features we know we want to make great, because we want that one to exist in the world, and also a lot of blue sky around what this could be. That requires an ability to go end to end in both design and engineering, but also be somewhat comfortable with the fact that it's experimental, meaning there will be a lot of uncertainty about whether this thing will actually ship. Because you want somebody to live with it, there's also a lot of work that needs to go into actually shipping it in some form so they can try.
蓝天空想的部分对我来说真的很有趣,因为即使只是听你讲了一会儿。
The blue sky piece is really interesting to me because even just listening to you talk now for a while.
我的意思是,很明显你正在做的事情与在 Mobin 上寻找并拼凑世界上已有的正确部分来找出解决方案完全相反,你真的是在最真实的意义上从第一性原理出发。那么我的问题是,你在候选人身上会寻找哪些信号,才能让你有信心说‘是的,他们确实有能力推动进展,帮助我们找到那些阶梯上的下一级’?
I mean it's so clear what you're doing is the complete opposite of looking on Mobin and trying to piece together the right pieces the right things that already exist in the world to like figure out a solution you know like you are really working from first principles in the truest sense. My question then is what are the signals that you would even look for in a candidate that would get you to a confidence level where you're like yeah they actually have what it takes to move the needle and help us find some of those next rungs on the ladder.
我们最优化的是使用界面时的实际感受。所以,尽管我们试图在原则性上确定一个我们无法亲身体验的空间的构建模块应该是什么,但我认为我们都有一个普遍的雄心,就是做出真正令人惊叹和鼓舞人心的东西。我们都间接地受到参与者的启发,我们确实希望他们使用这个神奇的新玩具时的感觉是美妙的。任何有明确愿望去做这些事情的人。在实践中,这意味着即使是那些你可能会删除的东西,也需要投入大量额外的工作来让它变得出色。这相当于其他公司在产品即将发布的一年内会进行的打磨和调优阶段。我们尽可能多地将这些投入到一个实验中,因为参与者会使用它,而这也是原型制作能力(即工程方面)如此关键的地方。技术栈非常广泛。我认为最紧密的互动在于机器学习方面和实际界面方面之间,因为机器学习方面最终(至少目前)需要找到一种方法来标注某些东西。所以,标注最终在某种程度上与他们在屏幕上看到的内容紧密相关。标注在多大程度上捕捉到了真实的意图,即他们当时真正的想法。模型的能力决定了界面的能力,而界面的功能又在某些方面强烈地偏向于模型的能力。所以这两者之间有一种非常直接的关系,这使得主要的瓶颈在于你是否能构建出端到端的东西来尝试这些想法,因为很多工作只是离线数据,即尝试收集数据然后离线运行分析,这对于界面方面的任何东西来说通常都是死胡同。这意味着你必须做出真正完全交互的东西来很好地验证这一点,并可能同时与建模方面的人合作。如果你有一个非常疯狂的新想法想要尝试,那会涉及很多工作。它涉及收集数据的任务,涉及最终的界面以便参与者能够控制,还涉及理想情况下让他们独立使用的东西,从而给你更丰富的信号来判断效果如何。所以,那些对做所有这些事情感到兴奋,并且对让每个实验尽可能出色感到兴奋的人,这样当我们遇到死胡同时,我们有信心那不仅仅是因为它只是一个 MVP。我通常认为 MVP 的心态有时会适得其反,因为你基本上做了绝对最低限度的事情,却只得到了一个相当平庸的信号。我们尝试的每一次射击,我们都尽可能做到最好,因为结果最终会决定我们前进的方向。
The thing we optimize most for is feeling when it comes to what it's actually like to use the interface. So, as much as we try to be principled about what the building block should be for a space that we cannot experience ourselves, I think there's just a general ambition to just make something that's like really awesome and inspiring, I think we're all yeah, obviously like indirectly so inspired by our participants and we do want what it feels like for them to use these this magical new toy to be amazing. anybody who has a clear desire to do those things. That's usually what that means in practice is even for something that you know that you're probably going to delete, there's like a lot of extra work that uh goes into just making that amazing. The equivalent of kind of like the polishing phase and the tuning phase that you'll have at a lot of other companies when something is actually, you know, going to ship in a calendar year. We try to put as much of that as we can into just an experiment because a person will use it and also this is where the prototyping ability like the engineering side of the role is so key. The stack here is like very wide. I think the the tightest interplay is really between the machine learning side of the work and the actual interface side of the work because the machine learning side ultimately some form of that at least right now is well we need to figure out a way to label something. So like the label is ultimately what uh on some level is very tied to what they see on screen. How well that label captures the real intent here like what they were really thinking at that time. What the models can do defines what the interface can do and what the interface does in some ways heavily biases what the models can do. So there's kind of like a very direct relationship between those two things especially that makes the primary bottleneck here just like can you build out something to try these things end to end because a lot of stuff is just the offline data meaning like trying to collect data and then just run your analysis offline generally is a dead end uh for anything on like the UI side. So it means you have to make something that is like actually fully interactive to actually validate this well and potentially also work with folks on the modeling side. If you have a really wild new idea you want to try that involves a lot. It involves a task to collect the data. It involves some final interface so that they can control it and again it involves ideally something that they can live with independently and give you a much richer signal on how well it's going to work. So people that yeah are excited about doing all of that and also are excited about making each individual experiment you run as amazing as it can be such that we have confidence that when we hit a dead end it's not just because it was for example like an MVP. I generally think that like the mentality of like an MVP can sometimes be quite self-defeating because you've basically done the absolute minimum you possibly could have done and gotten like a pretty mid signal on something. the shots that we try to shoot, we try to like shoot them as best we can because the the results that come out of that very much dictate ultimately the direction we're going to go in.
有没有一个例子,从一个直觉出发,引出了一系列实验过程,你可以从自己的经历中讲述一下,帮助人们更好地了解在 Neurolink 工作的人一天是什么样的?如果还能有一些例子,展示在整个实验中如何注重细节,那就更好了。
Is there an example of a hunch that led to some experimentation process that you could talk through from your own experience that would help people get a little bit better of a sense of what it is like day in the life of someone working on Neurolink? And maybe if there are also examples of what it looks like to sweat the details throughout that experiment, that would be great too.
一个很好的例子是我们目前称之为身体映射任务的任务。这个任务的结构是,我们渲染一个真实的 3D 手臂,并有一个完整的渲染管线来绘制这个 3D 手臂。这种方法意味着我们可以将你的注意力吸引到它的特定部分,并结合我们给你的实际指导,以及指导中的视觉指示器。那么这个任务是做什么的呢?这是我们在规格说明中的第一件事。这将是参与者坐下来做的第一件事。在他们真正进入校准植入物并移动光标的体验部分之前,他们只是探索一个广阔的空间:你现在要再次想象移动你的手臂。这只手臂你目前无法移动或只有非常有限的残余运动,他们感受一下什么对他们有效,哪种运动仍然感觉直观,即使他们不能很好地移动它。哪些不行?然后还要感受——不是感受,而是信号——哪些运动我们实际上能很好地读取。所以当他们想象这样做时,活动爆发是什么样的?哪些运动更好,哪些更弱,并在两者之间进行比较,从这一大组中选择一个,比如移动手臂上下左右,移动手腕上下左右,以及其他四种可能用于移动光标的方式。了解对他们感觉好的和我们实际能解码的运动之间的重叠。这个直觉主要是,在屏幕上看到那只手臂会让第一次体验比我们仅仅告诉他们“试试这样做”要自然得多。它会给他们一个可以观看和对照的东西。然后细节基本上是:好吧,我们该怎么做?我们可以尝试预渲染所有这些运动。那可能行得通。但如果我们能直接解码那只手臂呢?为什么要把这只手臂再降级到光标,如果他们已经在看那只手臂的世界里了?从某些方面来说,这比试图思考移动手臂来移动光标要自然得多。所以真正的直觉是,如果我们能做到这一点,他们是否也能基本上操控这只手臂?这需要大量的工程工作,因为我们不能再使用预渲染的动画了。我们需要实际渲染一个 3D 手臂。我们需要把它网格化。
One good example here is what we currently call the body mapping task. The structure of this task right now is there's an actual like 3D arm that we render and we have a whole rendering pipeline to draw this 3D arm. The approach there means that we can draw your focus to very specific pieces of it and then combine it with the actual guidance that we tell you with like visual indicators in the guidance that we give you as well. So what does this task do? This is the very first thing when we sort of spec this out. This would be the very first thing that participant sits down and does. And before they actually go into the part of the experience where they sort of calibrate their implant for the very first time and get to a moving cursor, they just explore this like wide space of like you are now going to imagine moving your arm again. This arm that like you currently cannot move or have very limited residual motion for and get a feel for just like what works for them like what kind of motion still feels intuitive to them even though they can't move it well. Which ones don't? and then also get a feel get a not a feel but a signal for what motions can we actually read really well. So when they imagine doing this, like what does the burst of activity look like? Um, and which motions seem to be better here and weaker and sort of compare between the two of like pick one from this wide set of like, you know, moving your arm up, down, left, right, moving your wrist up, down, left, right, and like four other things that you could presumably use to move a cursor. Getting some sense of the overlap between what feels good for them and also what uh we can actually decode. The hunch here was primarily that like seeing that arm on screening would make that first experience far more natural than us just telling them, hey, like try doing this. It would give them something to basically like look at and like cross reference. And then the details here were basically like, well, okay, well, how do we want to do this one? We could like try and pre-render all of these motions. Like that will probably work. But what if we could actually just decode that arm then? like why take this arm and then try to like go down to the cursor if they're already back in this world of just looking at that arm. In some ways that's a far more natural thing to do than trying to think about moving your arm to then move a cursor. So the real hunch here was if we can do this can we also like can they also puppet this arm basically to do that was an extraordinary amount of engineering work because we can't just use a pre-rendered animation anymore. We need to actually render like a 3D arm. We need to mesh that out.
我们需要一个着色管线,让手臂看起来不傻。屏幕上的 3D 手臂有巨大的恐怖谷效应,所以我们希望用最少的细节来传达手势,避免冗余信息。我们还希望它感觉自然,所以没有用骨架。最后,我们加了一点额外的魔法:当你做挤压动作时,文字会与手臂一起体现这个动作。我们做了动态排版,挤压文字会聚焦并随动作一起挤压。这完全没必要,但很酷。
We needed a shading pipeline so the arm wouldn't look goofy. There's a huge uncanny valley with a 3D arm on screen, so we wanted minimal detail to convey the gesture without redundant information. We also wanted it to feel natural, so we avoided a skeleton. Finally, we added an extra bit of magic: when you do a squeeze, text appears that embodies the motion alongside the arm. We did kinetic type where the squeeze text focuses in and squeezes with the motion. That was totally unnecessary but cool.
当时的直觉是,如果我们能通过屏幕上的手臂和正确的视角引发身心连接,用户就会想‘好了,我在挤压手臂’,而不是考虑关节或手指。然后我们就能解码这个动作。那将是一个丰富的交互空间——思考移动我的手臂,而不是光标。这条我再也无法移动的手臂。结果没成;解码整条手臂比解码光标差得多,这说得通,因为空间更大。但这是一个赌注:如果我们能解码整条手臂会怎样,以及如何将其与我们已有的目标统一起来。
The hunch was that if we could induce a mind-body connection with this arm on screen, with the right perspective, users would think, 'Okay, now I'm squeezing the arm,' not about joints or fingers. Then we could decode that. That would be a rich interaction space—thinking about moving my arm, not a cursor. This arm I can no longer move. It didn't pan out; decoding the full arm was far worse than decoding the cursor, which makes sense because it's a bigger space. But it was a bet on what if we could decode the full arm, and how to unify that with what we already want to do.
有趣的是,我读了你打算招的第二位设计师的职位描述。有一行我记在了笔记里:‘他们需要有雄心去构建超越局部最优解、让人感觉神奇的东西。’我觉得你用挤压文字的例子已经回答了这个问题。完全没必要,但带来愉悦的有趣方式。你在实验中做到了这一点,这完美诠释了即使不知道技术上是否可行也要超越期望。
It's funny, I read the job description for the second designer you're trying to bring on. There was a line I pasted into my notes: 'They need the ambition to build something that feels magical beyond the local minima of that which works.' I think you already answered that with the squeezing text. Totally unnecessary, but a fun way to bring delight. You did that in an experiment, which is the perfect example of going above and beyond even when you don't know if it's technically possible.
好了,Rooz,这期节目太棒了。感谢你创造了最独特的节目之一,可能是我做过的最独特的一期。听你思考交互和设计机会的方式真的非常迷人。非常感谢你今天来做客并分享。
Well, Rooz, this has been amazing. Thank you for creating one of the most unique episodes, probably the most unique episode I've ever had. Hearing the way you think about interactions and design opportunities is truly fascinating. Really appreciate you coming on and sharing with us today.
嗯,谢谢邀请。这非常有趣。
Yeah, thanks for having me. This was a ton of fun.
在结束之前,我想花一分钟介绍一下我最喜欢的产品,因为我经常被问到我的工具栈。Framer 用来建网站,Genway 用来做研究,Granola 用来在评审时记笔记,Jitter 用来做设计动画,Lovable 用来用代码实现想法,Mobin 用来找设计灵感,Paper 用来像创意人一样设计,Raycast 是我每一步的快捷方式。我精心挑选了这些公司,这样我才能全职做这些节目。所以,支持节目最好的方式就是去看看它们。你可以在 dive.comclub/partners 找到完整列表。
Before I let you go, I want to take just one minute to run you through my favorite products because I'm constantly asked what's in my stack. Framer is how I build websites. Genway is how I do research. Granola is how I take notes during crit. Jitter is how I animate my designs. Lovable is how I build my ideas in code. Mobin is how I find design inspiration. Paper is how I design like a creative. And Raycast is my shortcut every step of the way. Now, I've hand selected these companies so that I can do these episodes full-time. So, by far the number one way to support the show is to check them out. You can find the full list at dive.comclub/partners.