播客 · 访谈
#490 – State of AI in 2026: LLMs, Coding, Scaling Laws, China, Agents, GPUs, AGI
Nathan Lambert and Sebastian Raschka are machine learning researchers, engineers, and educators. Nathan is the post-training lead at the Allen Institute for AI (Ai2) and the author of The RLHF Book. Sebastian Raschka is the author of Build a Large Language Model (From Scratch) and Build a Reasoning Model (From Scratch). Thank you for listening ❤ Check out our sponsors: https://lexfridman.com/sponsors/ep490-sc See below for timestamps, transcript, and to give feedback, submit questions, contact Lex, etc. Transcript: https://lexfridman.com/ai-sota-2026-transcript CONTACT LEX: Feedback – give feedback to Lex: https://lexfridman.com/survey AMA – submit questions, videos or call-in: https://lexfridman.com/ama Hiring – join our team: https://lexfridman.com/hiring Other – other ways to get in touch: https://lexfridman.com/contact SPONSORS: To support this podcast, check out our sponsors & get discounts: Box: Intelligent content management platform. Go to https://box.com/ai Quo: Phone system (calls, texts, contacts) for businesses. Go to https://quo.com/lex UPLIFT Desk: Standing desks and office ergonomics. Go to https://upliftdesk.com/lex Fin: AI agent for customer service. Go to https://fin.ai/lex Shopify: Sell stuff online. Go to https://shopify.com/lex CodeRabbit: AI-powered code reviews. Go to https://coderabbit.ai/lex LMNT: Zero-sugar electrolyte drink mix. Go to https://drinkLMNT.com/lex Perplexity: AI-powered answer engine. Go to https://perplexity.ai/ OUTLINE: (00:00) – Introduction (01:39) – Sponsors, Comments, and Reflections (16:29) – China vs US: Who wins the AI race? (25:11) – ChatGPT vs Claude vs Gemini vs Grok: Who is winning? (36:11) – Best AI for coding (43:02) – Open Source vs Closed Source LLMs (54:41) – Transformers: Evolution of LLMs since 2019 (1:02:38) – AI Scaling Laws: Are they dead or still holding? (1:18:45) – How AI is trained: Pre-training, Mid-training, and Post-training (1:51:51) – Post-training explained: Exciting new research directions in LLMs (2:12:43) – Advice for beginners on how to get into AI development & research (2:35:36) – Work culture in AI (72+ hour weeks) (2:39:22) – Silicon Valley bubble (2:43:19) – Text diffusion models and other new research directions (2:49:01) – Tool use (2:53:17) – Continual learning (2:58:39) – Long context (3:04:54) – Robotics (3:14:04) – Timeline to AGI (3:21:20) – Will AI replace programmers? (3:39:51) – Is the dream of AGI dying? (3:46:40) – How AI will make money? (3:51:02) – Big acquisitions in 2026 (3:55:34) – Future of OpenAI, Anthropic, Google DeepMind, xAI, Meta (4:08:08) – Manhattan Project for AI (4:14:42) – Future of NVIDIA, GPUs, and AI compute clusters (4:22:48) – Future of human civilization
Lex Fridman Podcast · Lex Fridman
查单词
The following is a conversation all about the state-of-the-art in artificial intelligence,
接下来的对话全部关于人工智能领域的最新技术,
including some of the exciting technical breakthroughs and developments in AI that happened over the past year,
包括过去一年中人工智能领域发生的一些令人兴奋的技术突破和发展,
and some of the interesting things we think might happen this upcoming year.
以及我们认为今年可能发生的一些有趣的事情。
At times, it does get super technical,
有时候,内容确实会非常技术性,
but we do try to make sure that it remains accessible to folks outside the field without ever dumbing it down.
但我们确实努力确保它对领域外的人也能理解,同时绝不降低深度。
dumbing it down 常用搭配
把内容简化到降低其深度或水准
常带否定,表示在让内容易懂的同时不牺牲专业性。
It is a great honor and pleasure to be able to do this kind of episode with two of my favorite people in the AI community, Sebastian Rashka and Nathan Lambert.
能够与人工智能社区中我最喜欢的两个人——塞巴斯蒂安·拉什卡和内森·兰伯特——一起做这样一期节目,是我的莫大荣幸和快乐。
They are both widely respected machine learning researchers and engineers who also happen to be great communicators, educators, writers, and Twitterers, ex-posters.
他们都是广受尊敬的机器学习研究者和工程师,同时也恰好是出色的沟通者、教育者、作家和推特用户,前发帖人。
Sebastian is the author of two books I highly recommend for beginners and experts alike.
塞巴斯蒂安是两本书的作者,我强烈推荐给初学者和专家。
First is build a large language model from scratch and build a reasoning model from scratch.
第一本是《从零构建大型语言模型》和《从零构建推理模型》。
I truly believe in the machine learning computer science world, the best way to learn and understand something is to build it yourself from scratch.
我真心相信,在机器学习计算机科学领域,学习和理解某件事的最佳方式就是自己从零开始构建它。
from scratch 常用搭配
从零开始,从头做起
表示不借助现成基础,完全从头构建或学习某事物。
Nathan is the post-training lead at the Allen Institute for AI and author of the definitive book on reinforcement learning from human feedback.
内森是艾伦人工智能研究所的后训练负责人,也是关于从人类反馈中进行强化学习的权威书籍的作者。
Both of them have great X accounts, great substacks.
他们两人都有很棒的X账号和很棒的Substack。
Sebastian has courses on YouTube.
塞巴斯蒂安在YouTube上有课程。
Nathan has a podcast.
内森有一个播客。
And everyone should absolutely follow all of those.
大家都绝对应该关注所有这些。
And now a quick few second mention of each sponsor.
现在快速用几秒钟提一下每个赞助商。
Check them out in the description or at lexfriedman.com/sponsors.
请在描述中或访问 lexfriedman.com/sponsors 查看它们。
It is, in fact, the best way to support this podcast.
事实上,这是支持这个播客的最佳方式。
We got a bunch of great sponsors.
我们有很多很棒的赞助商。
for intelligent content management, Quo for a phone system
用于智能内容管理,Quo 用于电话系统
like call stacks, contacts for your business,
比如呼叫栈、企业联系人,
Uplift Desk, the desk I'm sitting behind, and my favorite office desk.
Uplift Desk,我正坐在后面的桌子,也是我最喜欢的办公桌。
Fin for customer service AI agents,
Fin 用于客户服务 AI 代理,
Shopify for selling stuff online,
Shopify 用于在线销售,
CodeRabbit for AI-powered code review,
CodeRabbit 用于 AI 驱动的代码审查,
Element for electrolytes,
Element 用于电解质,
and of course, our longtime friend, Perplexity, for curiosity-driven knowledge exploration.
当然,还有我们的老朋友 Perplexity,用于好奇心驱动的知识探索。
Choose wisely, my friends.
明智地选择吧,我的朋友们。
And now, on to the full ad reads.
现在,进入完整的广告朗读。
I try to make them interesting, but if you do skip,
我尽量让它们有趣,但如果你跳过,
please still check out the sponsors.
请仍然查看赞助商。
I enjoy their stuff.
我喜欢他们的东西。
Maybe you will, too.
也许你也会喜欢。
To get in touch with me, for whatever reason, go to lexfreemer.com/contact.
要联系我,无论什么原因,请访问 lexfreemer.com/contact。
If you can't tell, I'm trying to have a bit of a pep in my step.
如果你看不出来,我正试着让自己精神一点。
a pep in my step 常用搭配
步伐轻快、精神抖擞的样子
口语中形容人显得有活力、心情振奋。
At the moment.
目前。
Because I had a long night, didn't get much sleep at all.
因为我熬了一夜,根本没怎么睡。
So I am running on fumes, delirious, happy, unsure of what is reality and what is a dream.
所以我靠着一股劲在撑,神志不清,开心,分不清什么是现实什么是梦境。
running on fumes 常用搭配
精疲力竭却还在硬撑
口语中形容极度疲劳、靠最后一点精力维持。
In fact, we could right now be living inside of a dream.
事实上,我们现在可能正生活在一个梦里。
I have been going through a lot.
我经历了很多。
I have been working insane hours, so much going on.
我一直都在疯狂地工作,有太多事情在发生。
I am so overwhelmed.
我真是太不堪重负了。
Of course, as always, truly grateful and happy to be alive.
当然,一如既往,真心感激并且高兴自己还活着。
but have not been able to publish as many episodes as you would like,
但没能发布像你希望的那样多的节目,
so there's a bunch of sponsors we have to catch up on.
所以有一堆赞助商我们需要补上。
Your support truly means the world.
你的支持真的意义重大。
means the world 常用搭配
意义极其重大,非常重要
表达感激或情感时用,表示某事物对自己无比重要。
Please check out all the sponsors.
请查看所有的赞助商。
If you think it might be useful to you, buy their stuff.
如果你觉得可能对你有用,就买他们的东西。
It really is the best way to support this podcast.
这真的是支持这个播客的最好方式。
All right, let's go.
好的,我们开始吧。
First up, this episode is brought to you by Box,
首先,本期节目由 Box 赞助,
a cloud-based platform for content management, file sharing,
一个基于云的内容管理、文件共享平台,
And all kinds of collaboration, all kinds of content for your businesses.
以及各种协作,各种适合你企业的内容。
Like with a lot of companies, the big question is,
就像很多公司一样,最大的问题是,
how is AI leveraged to make whatever the business does better?
如何利用 AI 让企业所做的任何事情变得更好?
A lot of companies kind of use it for the hype and the label.
很多公司有点把它用于炒作和标签。
It's kind of hilarious to watch people just say like, powered by AI.
看着人们只是说“由 AI 驱动”,有点好笑。
I don't care if you're a bakery, powered by AI.
我不在乎你是不是一家面包店,由 AI 驱动。
I don't know.
我不知道。
But outside of all of the hype,
但抛开所有的炒作,
it is one of the most incredible things that humans have ever created.
它是人类创造过的最不可思议的事物之一。
And so companies that can leverage that well are the companies that win.
所以能够很好地利用它的公司就是赢家。
And of course, Box is legendary for its file and content management,
当然,Box 因其文件和内容管理而闻名,
especially when you're talking about scale.
尤其是在谈到规模的时候。
So obviously, it's amenable for the utilization of AI to help automate some of the document processing, some of the workflow, some of the organization.
所以很明显,它适合利用AI来帮助自动化一些文档处理、一些工作流程、一些组织工作。
and they do that exceptionally well.
而且他们做得非常出色。
They have a system called, as you could imagine, Box AI that does just that.
他们有一个系统,正如你可以想象的,叫做Box AI,正是做这个的。
I love it.
我很喜欢它。
They do an excellent implementation on the interface side, on the backend side.
他们在界面端和后端都做了出色的实现。
Everything works extremely nicely.
一切都运行得非常好。
Help scale AI across your organization today and go to box.com slash AI.
今天就帮助在你的组织中扩展AI,请访问box.com/AI。
That's box.com slash AI to learn more.
那就是box.com/AI,了解更多。
This episode is also brought to you by Quo, spelled Q-U-O.
本期节目还由Quo赞助,拼写为Q-U-O。
Also happens to be a company name with just three letters that will help you win at Scrabble.
它恰好也是一个只有三个字母的公司名,能帮你在拼字游戏中获胜。
Are you allowed to use company names with Scrabble?
拼字游戏允许使用公司名称吗?
How many points is Q?
Q值多少分?
How many points is U?
U值多少分?
I'm imagining a lot.
我想象有很多分。
That was one of the big confusions to me when I was first learning the English language.
这是我最初学习英语时最大的困惑之一。
It always felt like Q should be at the end of the alphabet, maybe like QZ.
总觉得Q应该在字母表的末尾,也许像QZ那样。
It was always surprising to my limited brain capacity that Q was earlier on in the alphabet.
以我有限的大脑容量,Q在字母表中更靠前总是让我感到惊讶。
What is it?
它是什么?
OPQ?
OPQ?
I can't even actually localize letters in the alphabet.
我甚至无法在字母表中定位字母。
I'm sure that's the case for a lot of people without reading the alphabet in my head sequentially.
我相信很多人都是这样,除非我在脑子里按顺序读字母表。
that's the case for 句型
……的情况也是如此
that's the case for [someone/something]
用于表示某种情况对某人或某群体同样适用,常与 for 搭配引出对象。
All of this has to do with short-term and long-term memory access,
所有这些都与短期和长期记忆的获取有关,
has to do with 常用搭配
与……有关
口语中常用来说明事物之间的关联,比 be related to 更自然随意。
the functioning, the limitation of human cognition,
人类认知的运作和局限,
and maybe cognitive systems in general,
也许还有一般的认知系统,
all of it relevant to this particular episode,
所有这些都与这一集特别相关,
and not so relevant to the awesomeness of Quo, formerly known as OpenPhone,
而与Quo(原名OpenPhone)的精彩之处关系不大,
formerly known as 常用搭配
原名,以前被称为
用于介绍某人或某物改名前的旧名称,常见于正式或介绍性语境。
that I should be talking about.
那才是我应该谈论的。
Of course, as is always the case,
当然,一如既往,
as is always the case 常用搭配
一如既往,情况总是如此
用于引出通常会发生或普遍存在的情况,语气略带书面或正式。
I think the point here and at the point everywhere in the point of life
我认为这里的重点,以及生活中处处都是重点的重点,
is to talk from the heart about whatever you want.
就是发自内心地谈论任何你想谈的事。
from the heart 常用搭配
发自内心地
形容说话或做事真诚、不虚伪,常用于表达真实情感。
And that's what I try to do with everything.
这就是我对待一切事情的方式。
And to generalize that even more, to talk whenever I want and to shut the F up
再进一步概括来说,就是想说的时候就说,不想说的时候就闭嘴,
whenever I want and listen.
想听的时候就听。
And I prefer that more often than I prefer to talk.
而且比起说话,我更常喜欢这样。
Insert clever transition here because talk is somehow relevant, it is.
在这里插入一个巧妙的过渡,因为谈话不知怎么地确实相关,确实如此。
So Quo, formerly known as OpenPhone, helps over 90,000 businesses manage phone calls, texts, contacts,
所以Quo,原名OpenPhone,帮助超过9万家企业管理电话、短信、联系人,
all kinds of phone-related stuff for a business.
以及各种与企业电话相关的事务。
You have a bunch of customers, a bunch of incoming calls,
你有一大堆客户,一大堆来电,
a bunch of 常用搭配
一大堆,许多
非正式口语,表示数量多,比 a lot of 更随意。
a bunch of people on the business side that have to answer those calls,
企业这边有一大堆人必须接听这些电话,
have to manage it, what's the status of this particular request,
必须管理这些,这个特定请求的状态是什么,
voicemails, transcripts, all that kind of stuff,
语音留言、文字记录,诸如此类,
all that kind of stuff 常用搭配
诸如此类的东西
口语中用于列举后收尾,表示还有其他类似的事物,不必一一列出。
and obviously a really nice, effective utilization of AI to make that really efficient.
显然,这是对AI非常好、非常有效的利用,让这一切变得非常高效。
But really, what's really important for things like this is that the interface is good,
但真正重要的是,对于这类东西来说,界面要好,
that team collaboration is good, and Quo delivers on that.
团队协作要好,而Quo在这方面做到了。
delivers on that 常用搭配
在这一点上做到了,兑现了承诺
用于表示某人或某产品达到了预期或承诺的标准,常与 promise 等词搭配。
Try Quo for free, plus get 20% off your first six months when you go to quo.com slash lex.
免费试用Quo,另外访问quo.com/lex,前六个月还可享受20%折扣。
That's Q-U-O dot com slash lex.
网址是Q-U-O点com斜杠lex。
Tell your friends about it, because it just might help them win at Scrabble.
告诉你的朋友们,因为这可能正好帮他们在拼字游戏中获胜。
Speaking of Scrabble, you usually want to play Scrabble on a table.
说到拼字游戏,你通常想在桌子上玩拼字游戏。
Speaking of 地道口语
说到,提到
用于自然地引出与刚提到的话题相关的新内容,是口语中常见的过渡语。
It's such a magical experience.
这是一种如此神奇的体验。
I just had a vision from a distant past of me sitting with a friend and playing Scrabble at a table.
我刚刚浮现出一个遥远过去的画面,我和一个朋友坐在桌边玩拼字游戏。
What is this life full of beautiful memories and that it's over too soon?
这是什么人生,充满了美好的回忆,却结束得太快?
Yeah. That melancholy feeling is beautiful, I think.
是的。我觉得那种忧郁的感觉很美。
Insert another clever transition, a la Mark Norman maybe, because the name of this next company is Uplift Desk.
插入另一个巧妙的过渡,也许像马克·诺曼那样,因为下一家公司的名字是Uplift Desk。
As I said. Okay, it's my go-to favorite office desk, and it's also the desk that I use for podcast furniture.
正如我所说。好的,这是我最喜欢、最常用的办公桌,也是我用来做播客家具的桌子。
go-to 常用搭配
首选,最常用的
形容某物是遇到某种需求时首先想到或最常使用的选择。
I have, I already lost count, I have a lot of Uplift desks, standing desks in my place, everywhere.
我已经数不清了,我有很多Uplift桌子,站立式桌子,到处都是。
lost count 常用搭配
数不清了
表示数量太多以至于无法计数,常用于口语中强调数量多。
It's desks everywhere. I have a mattress on the floor and Uplift desks.
到处都是桌子。我有一张床垫放在地板上,还有Uplift桌子。
So I have a Linux box for robotics.
所以我有一台用于机器人技术的Linux主机。
I have a machine where I do a lot of the editing.
我有一台机器,我在上面做很多剪辑工作。
All of that is on a desk.
所有这些都放在桌子上。
I have the three tables for the podcast desk,
我有三张用于播客桌的桌子,
the very one you've seen over the past several years.
就是你过去几年一直看到的那一张。
That's all uplift desks.
那都是Uplift的桌子。
I usually don't put them in standing mode,
我通常不把它们调到站立模式,
but they are standing desks.
但它们是站立式桌子。
It allows me to do all kinds of stuff,
它让我能做各种事情,
really easy to work with,
真的很容易使用,
really nice material,
材质真的很好,
really sturdy.
真的很结实。
I just love everything about uplift desk.
我就是喜欢Uplift桌子的一切。
When they said they want to sponsor,
当他们说想要赞助时,
After I've been using them for many years,
在我用了它们很多年之后,
I lost my mind.
我简直疯了。
lost my mind 地道口语
兴奋得发疯,激动坏了
口语中夸张地表达极度兴奋或震惊,不一定是真的失去理智。
I love it when I've been in love with a company,
我喜欢当我爱上一家公司的时候,
in love with 常用搭配
爱上,非常喜欢
口语中可夸张地表示对某人、某物或某公司极其喜爱。
in love with their product for such a long time,
爱上他们的产品这么长时间,
and I get to also sing them praises.
而且我还能赞美他们。
sing them praises 常用搭配
赞美他们,高度赞扬
表示公开或热情地称赞某人或某物,常用 sing someone's praises 形式。
I mean, come on.
我是说,拜托。
I mean, come on 地道口语
我是说,拜托
口语中用于表达难以置信、无奈或强调显而易见的事,语气随意。
What are you going to tell me next that FFMPEG wants to sponsor this podcast?
接下来你还要告诉我FFMPEG想赞助这个播客吗?
Another sort of open source project.
另一个开源项目。
It's not a company that I've been in love with.
这不是一个我一直热爱的公司。
Anyway, go to upliftdesk.com slash Lex and use code Lex to get four free accessories
总之,去 upliftdesk.com/lex 并使用代码 Lex 获得四个免费配件,
free same-day shipping,
免费当日送达,
free returns,
免费退货,
a 15-year warranty,
15年保修,
and an extra discount off your entire order.
以及整个订单的额外折扣。
That's upliftdesk.com.
那就是 upliftdesk.com。
The spelling it out really help anybody.
把它拼出来真的能帮到任何人。
I don't know,
我不知道,
but they really said pretty please.
但他们真的说了拜托了。
pretty please 地道口语
拜托了,求求你了
口语中用于撒娇或半开玩笑地请求,比 please 更强调恳求。
The one request is spell it out.
唯一的要求就是把它拼出来。
Again, what is this life?
再说一次,这是什么生活?
Incredible.
不可思议。
This episode is also brought to you by Finn, the number one AI agent for customer service.
本期节目同样由 Finn 赞助,它是客户服务领域的头号 AI 代理。
Find the niche and become number one.
找到细分市场,成为第一名。
That's the idea here.
这就是这里的理念。
Anybody building an AI company, and we talk about this, is the dream of AGI dead?
任何正在创建 AI 公司的人,我们谈论这个话题,AGI 的梦想死了吗?
I think for a lot of companies, success is in the niche,
我认为对很多公司来说,成功在于细分市场,
but there is a few, and Finn delivers on that niche.
但有一些公司,Finn 在那个细分市场做到了。
delivers on that niche 常用搭配
在那个细分市场做到了、兑现了承诺
用于描述公司或个人在某个特定细分领域真正做出成绩、满足需求。
It's trusted by over 6,000 customer service leaders at top companies, including AI companies.
它受到超过 6,000 名顶级公司客户服务负责人的信任,包括 AI 公司。
When an AI company trusts your company to do its customer service, that means you're legit.
当一家 AI 公司信任你的公司来做它的客户服务时,这意味着你是靠谱的。
you're legit 地道口语
你是靠谱的、是正规可信的
口语中表示某人或某事物真实可信、有实力,非正式场合常用。
90-day money bag guarantee up to $1 million built to handle complex multi-step priorities like returns, exchanges, and disputes.
90 天资金保证,最高 100 万美元,专为处理复杂的多步骤优先事项而构建,如退货、换货和纠纷。
Go to fin.ai.lex to learn more about transforming your customer service and scaling your support team.
访问 fin.ai.lex 了解更多关于转变您的客户服务和扩展您的支持团队的信息。
That's fin.ai slash lex.
那就是 fin.ai/lex。
I don't know why I switched to this hyping voice.
我不知道为什么我切换到了这种夸张的声音。
Crappy announcer, crappy radio jockey, crappy ad read voice.
糟糕的播音员,糟糕的电台主持人,糟糕的广告朗读声音。
It is what it is.
事情就是这样。
It is what it is. 地道口语
事情就是这样,没办法。
口语中表示接受无法改变的现实,带一点无奈。
Thank you for sticking with me this long.
感谢你陪我这么久。
sticking with me 常用搭配
一直陪着我、坚持听/看下去
用于感谢听众或观众长时间陪伴、没有中途离开。
I feel the love and I send it right back at you.
我感受到了爱,我也把爱回送给你。
I send it right back at you 地道口语
我也把同样的(爱)回送给你
口语中回应别人的情感表达,表示同样的感情回赠对方。
This episode is also brought to you by a company whose engineers are also full of love, Shopify.
本期节目同样由一家工程师也充满爱的公司赞助,Shopify。
It just brings a smile to my face.
它只是让我脸上露出微笑。
brings a smile to my face 常用搭配
让我脸上露出微笑、让我开心
用于表达某事物让人感到愉快、温暖。
Every time I think about Shopify, I got to see their engineering booth at NeurIPS, which is a machine learning conference.
每次我想到Shopify,我都会想起在NeurIPS看到他们的工程展台,那是一个机器学习会议。
Really brilliant people, wonderful people.
非常聪明的人,很棒的人。
Of course, the CEO, Toby, is still programming, still building stuff, still in on the details of the engineering,
当然,CEO Toby仍然在编程,仍然在构建东西,仍然深入工程的细节,
in on the details 常用搭配
深入参与、了解细节
用于形容领导者或高层仍然亲自关注具体细节。
and now is talking quite a bit about utilization of LLMs for his own sort of pet projects, but also inside the company.
而现在他经常谈论LLM的利用,既用于他自己的个人项目,也用于公司内部。
It's just incredible when, from the very top, the company is in love with engineering.
当公司从最高层就热爱工程时,这真是令人难以置信。
It's a celebration of great engineering,
这是对伟大工程的庆祝,
Just like the conversation with DHH, who is the guy behind Ruby on Rails that Shopify was built on.
就像与DHH的对话一样,他是Shopify所基于的Ruby on Rails背后的人。
That conversation was a celebration of great engineering.
那次对话是对伟大工程的庆祝。
The beauty of engineering as well.
也是工程之美。
Anyway, listen to that episode to see some of the magic of Ruby on Rails and the magic of Shopify and the magic of Toby that we talk about.
总之,听听那一集,看看我们谈论的Ruby on Rails的魔力、Shopify的魔力以及Toby的魔力。
Anyway, sign up for a $1 per month trial period at shopify.com slash lux. That's all lowercase.
总之,在shopify.com/lux注册每月1美元的试用期。全部小写。
Go to shopify.com slash lex to take your business to the next level today.
访问shopify.com/lex,今天就让你的业务更上一层楼。
take your business to the next level 常用搭配
让你的业务更上一层楼
商业广告或建议中常用,表示帮助业务提升到更高水平。
This episode is also brought to you by CodeRabbit, a platform that provides AI-powered code reviews directly within your terminal.
本期节目还由CodeRabbit赞助,这是一个直接在终端中提供AI驱动代码审查的平台。
We talk a lot in this episode about the timeline for the full automation of the human programmer.
我们在本期节目中大量讨论了人类程序员完全自动化的时间线。
I think we're quite far away from taking the human out of the loop.
我认为我们离把人类排除在循环之外还差得很远。
taking the human out of the loop 常用搭配
把人类排除在流程之外
用于讨论自动化时,指不再需要人类参与某个环节。
That review process, the debugging process, all of that, that's such a crucial part of programming,
那个审查过程、调试过程,所有这些,都是编程中极其关键的一部分,
especially just like we talk about in the episode.
尤其是就像我们在这一集里谈到的那样。
When we're not talking about a personal website where HTML slop is something that a web browser magically, automagically,
当我们不是在谈论一个个人网站,那里的HTML垃圾是网页浏览器神奇地、自动地
I don't know how they're possibly able to do such incredible job of rendering slop,
我不知道它们怎么可能把垃圾渲染得如此出色,
but a web browser is in fact able to render slop, including AI slop.
但网页浏览器实际上确实能够渲染垃圾,包括AI垃圾。
It just finds a way. So really the question is when you have production code, something that a lot of users are relying on,
它总能找到办法。所以真正的问题是,当你有生产代码,也就是很多用户依赖的东西时,
It just finds a way. 地道口语
它总能找到办法。
口语中形容某事物总能自行解决问题或达成结果。
how do you review that code? How do you make sure you're catching errors?
你如何审查那段代码?你如何确保能捕捉到错误?
How are you making sure that you put a backstop to hallucinations and the logical errors that AI coding agents can generate.
你如何确保为幻觉以及AI编码代理可能产生的逻辑错误设置一道防线。
put a backstop to 常用搭配
为……设置一道防线
用于描述为防止错误或风险而设置的防护措施。
Anyway, CodeRabbit supports all programming languages.
总之,CodeRabbit支持所有编程语言。
Install CodeRabbit CLI today at coderabbit.ai.lex.
今天就在coderabbit.ai.lex安装CodeRabbit CLI。
That's coderabbit.ai.lex.
就是coderabbit.ai.lex。
This episode is also brought to you by Element,
本期节目还由Element赞助,
my daily zero-sugar and delicious electrolyte mix.
我每天喝的无糖又美味的电解质冲剂。
Reminds me of the fact that I need to get to editing the video of me in the jungle when
这让我想起,我得去剪辑我在丛林里的视频,当时
Paul Rosling and I, such an incredible human,
保罗·罗斯林和我,多么了不起的人,
congratulations to Paul on all of his success.
祝贺保罗取得的所有成功。
Go get his book. It's an incredible book.
去买他的书吧。那是一本了不起的书。
Again, he's an incredible person with an incredible mission.
再说一次,他是一个了不起的人,肩负着了不起的使命。
And yes, I need to edit and publish, hoping to at the very least.
是的,我需要编辑并发布,至少希望能做到这一点。
at the very least 常用搭配
至少,最起码
用于表示最低限度的期望或让步。
The story of our journey in the jungle,
我们在丛林中的旅程故事,
because it was a beautiful celebration of nature and the jungle and friendship and the full richness of the human experience.
因为那是对自然、丛林、友谊以及人类经验全部丰富性的一次美丽庆祝。
It was beautiful.
那很美。
The reason I mentioned that is I was, as part of that journey, severely dehydrated.
我提到这个的原因是在那段旅程中,我严重脱水。
And I remember dreaming of element of a cold drink of water with the electrolytes.
我记得我梦见元素,一杯带电解质的冷饮。
Your body craves it.
你的身体渴望它。
And it craves it because it needs it.
它渴望它,因为它需要它。
Electrolytes, sodium, potassium, magnesium.
电解质、钠、钾、镁。
When you're deprived, it's not just water, it's electrolytes.
当你缺乏时,不仅仅是水,还有电解质。
So anyway, I always remember that.
总之,我一直记得这一点。
Get a free eight-count sample pack with any purchase.
任何购买都可获得免费八支装样品包。
Try it at drinkelement.com slash lex.
请在 drinkelement.com/lex 试用。
This is the Lex Friedman Podcast.
这是莱克斯·弗里德曼播客。
To support it, please check out our sponsors in the description, where you can also find links to contact me, ask questions, get feedback, and so on.
为了支持它,请查看描述中的赞助商,在那里你还可以找到联系我、提问、获取反馈等的链接。
And now, dear friends, here's Sebastian Rashka and Nathan Lambert.
现在,亲爱的朋友们,有请塞巴斯蒂安·拉什卡和内森·兰伯特。
So I think one useful lens to look at all of this through is the DeepSeq, so-called DeepSeq moment.
所以我认为看待这一切的一个有用视角是 DeepSeq,所谓的 DeepSeq 时刻。
This happened about a year ago in January 2025
这大约发生在一年前,2025年1月
when the open-weight Chinese company DeepSeq released DeepSeq R1
当时中国开源权重公司DeepSeq发布了DeepSeq R1
that I think it's fair to say surprised everyone with near or at state-of-the-art performance
我认为可以说,它以接近甚至达到最先进的性能让所有人惊讶
it's fair to say 常用搭配
可以说,公平地说
用于表达一个较为客观、被普遍认可的判断。
with allegedly much less compute for much cheaper.
据称计算量少得多,成本也便宜得多。
and from then to today, the AI competition has gotten insane,
而从那时到今天,AI竞争已经变得疯狂,
has gotten insane 地道口语
已经变得疯狂、极其激烈
口语中形容某事物程度极高、非常夸张。
both on the research level and the product level.
无论是在研究层面还是产品层面。
It's just been accelerating.
它一直在加速。
Let's discuss all of this today and maybe let's start with some spicy questions if we can.
我们今天来讨论这一切,也许我们可以从一些辛辣的问题开始。
Who is winning at the international level?
在国际层面上,谁在赢?
Would you say it's the set of companies in China or the set of companies in the United States?
你会说是中国的公司,还是美国的公司?
And Sebastian, Nathan, it's good to see you guys.
还有Sebastian、Nathan,很高兴见到你们。
So Sebastian, who do you think is winning?
那么Sebastian,你认为谁在赢?
So winning is a very broad term.
所以“赢”是一个非常宽泛的词。
I would say you mentioned the DeepSeek moment,
我想说,你提到了DeepSeek时刻,
and I do think DeepSeek is definitely winning the hearts of the people who work on open weight models
我确实认为DeepSeek绝对赢得了从事开源权重模型工作的人们的心
winning the hearts of 常用搭配
赢得……的心
用于描述获得某群体的喜爱和支持。
because they share these as open models.
因为他们把这些作为开放模型分享出来。
Winning, I think, has multiple timescales to it.
我认为,“赢”有多个时间尺度。
We have today, we have next year, we have in 10 years.
我们有今天,有明年,有10年后。
One thing I know for sure is that I don't think nowadays, 2026,
我确定知道的一件事是,我觉得如今,2026年,
I know for sure 常用搭配
我确信
用于强调自己对某事非常确定,口语中常用来引出有把握的判断。
that there will be any company who is, let's say, having access to a technology that no other company has access to.
会有任何一家公司,比如说,拥有其他公司无法获得的某项技术。
having access to 常用搭配
能够使用或获得
表示拥有使用某物(如技术、资源、信息)的途径或权利。
And that is mainly because researchers are frequently changing jobs, changing labs, they rotate.
这主要是因为研究人员经常换工作、换实验室,他们会轮换。
So I don't think there will be a clear winner in terms of technology access.
所以我认为在技术获取方面不会有明显的赢家。
in terms of 常用搭配
在……方面
用于限定讨论的范围或角度,说明从哪个方面来看。
However, I do think there will be the differentiating factor will be budget and hardware constraint.
不过,我确实认为差异化的因素将是预算和硬件限制。
So I don't think the ideas will be proprietary, but the way or the resources that are needed to implement them.
所以我认为想法不会是专有的,而是实现它们所需的方式或资源。
And so I don't see currently a take it all scenario where a winner takes it all.
所以目前我看不到一个赢家通吃的局面。
a winner takes it all 地道口语
赢家通吃
形容竞争中胜者拿走全部利益、没有其他人份额的局面。
I can't see that at the moment.
目前我看不到这种情况。
Nathan, what do you think?
内森,你怎么看?
You see the labs put different energy into what they're trying to do.
你可以看到各个实验室在他们想做的事情上投入了不同的精力。
put different energy into 常用搭配
在……上投入不同的精力
表示不同的人或团队对各自目标投入的努力程度或方式不同。
And I think to demarcate the point in time when we're recording this,
我想为了标明我们录制这个的时间点,
the hype over Anthropic Cloud Opus 4.5 model has been absolutely insane,
关于Anthropic Cloud Opus 4.5模型的炒作简直疯狂,
which is just I mean, I've used it and built stuff in the last few weeks.
我的意思是,我在过去几周里用过它并做了一些东西。
And it's it's almost gotten to the point where it feels like a bit of a meme in terms of the hype.
它几乎已经到了在炒作方面感觉有点像梗的地步。
gotten to the point where 句型
已经到了……的地步
get to the point where [clause]
用于描述某事发展到某个程度,后面接结果或状态。
And it's kind of funny because this is very organic.
这有点好笑,因为这是非常自然的。
And then if we go back a few months ago, we can get the release date in the notes as Gemini 3 from Google got released.
然后如果我们回到几个月前,我们可以在笔记里看到发布日期,是谷歌的 Gemini 3 发布了。
And it seemed like the marketing and just like wow factor of that release was super high.
而且那次发布的营销和那种哇塞效应似乎超级高。
But then at the end of November, Claude Opus 4.5 was released and the hype has been growing.
但到了11月底,Claude Opus 4.5 发布了,热度一直在增长。
But Gemini 3 was before this.
但 Gemini 3 是在这之前。
And it kind of feels like people don't really talk about it as much, even though when it came out, everybody was like, this is Gemini's moment to retake kind of Google's structural advantages in AI.
而且感觉人们并没有那么频繁地谈论它,尽管它刚出来时,大家都说,这是 Gemini 重新夺回谷歌在 AI 领域结构性优势的时刻。
people don't really talk about it as much 句型
人们不怎么像以前那样谈论它了
[people] don't really talk about [something] as much
用于表示某事物不再像之前那样受到关注或讨论。
And Gemini 3 is a fantastic model and I still use it.
而且 Gemini 3 是一个很棒的模型,我还在用它。
It's with sebastian what you're saying with all these like the idea space is very fluid but um culturally anthropic is known for betting very hard on code which is cloud code thing is working out for them right now so i think that even if the ideas flow pretty freely so much of this is bottlenecked by human effort and kind of culture of organizations where anthropic seems to at least be presenting as the least chaotic it's a bit of an advantage and if they can keep doing that for a while but
塞巴斯蒂安,你说的这些,就像整个创意空间非常流动,但嗯,在文化上 Anthropic 以在代码上押重注而闻名,也就是 Cloud Code 这件事现在对他们很有效,所以我认为即使想法流动得相当自由,这其中很多都受限于人力以及组织文化,而 Anthropic 似乎至少表现得是最不混乱的,这算是一个优势,如果他们能再保持一段时间,但
betting very hard on 常用搭配
在……上下重注
表示对某方向投入极大资源或信心,愿意承担风险。
working out for them 常用搭配
对他们来说进展顺利
表示某个做法或决定产生了好的结果,口语中常用。
On the other side of things, there's a lot of ominous technology from China,
另一方面,中国有很多令人不安的技术,
where there's way more labs than DeepSeek.
那里的实验室远不止DeepSeek一家。
So DeepSeek kicked off a movement within China.
所以DeepSeek在中国国内掀起了一场运动。
kicked off a movement 常用搭配
掀起了一场运动
表示某个事件引发了一股广泛的潮流或行动。
I'd say kind of similar to how ChatGPT kicked off a movement in the U.S.,
我会说这有点像ChatGPT在美国掀起的那场运动,
where everything had a chatbot.
那时所有东西都配了个聊天机器人。
There's now tons of tech companies in China that are releasing very strong frontier open weight models,
现在中国有大量科技公司正在发布非常强大的前沿开放权重模型,
to the point where I would say that DeepSeek is kind of losing its crown
以至于我会说DeepSeek正在失去它的王冠,
losing its crown 常用搭配
失去领先地位
比喻某个原本领先的人或事物正在失去优势或王座地位。
as the preeminent open model maker in China, and the likes of Z.ai's with their GLM models, Minimax's models, Kimi Moonshot, especially in the last few months, have shown more brightly.
作为中国最杰出的开放模型制造商,而像Z.ai的GLM模型、Minimax的模型、Kimi月之暗面,尤其是最近几个月,表现得更加亮眼。
The new DeepSeek models are still very strong,
新的DeepSeek模型仍然非常强大,
but that's kind of a, it could look back as a big narrative point
但这有点像,回头看可能是一个重大的叙事节点,
where in 2025 DeepSeek came and then all, and it kind of provided this platform for way more Chinese companies that are releasing these fantastic models to kind of have this new type of operation.
在2025年DeepSeek出现之后,它某种程度上为更多中国公司提供了一个平台,让它们发布这些出色的模型,从而拥有这种新型的运营模式。
So these models from these Chinese companies are open weights.
所以这些中国公司的模型是开放权重的。
And depending on the trajectory of business models that these American companies are doing could be at risk.
而取决于这些美国公司所采用的商业模式的走向,它们可能会面临风险。
at risk 常用搭配
面临风险
表示可能受到损害或处于不利境地。
But currently, a lot of people are paying for AI software in the U.S. and historically in China and other parts of the world,
但目前,很多人在美国为AI软件付费,而历史上在中国和世界其他地区,
people don't pay a lot for software.
人们不会为软件付很多钱。
So some of these models like DeepSeek have the love of the people because they are open weight.
所以像DeepSeek这样的一些模型受到人们的喜爱,因为它们是开放权重的。
the love of the people 常用搭配
民众的喜爱
表示受到大众欢迎或喜爱,常用于描述产品或人物。
How long do you think the Chinese companies keep releasing open weight models?
你认为中国公司会继续发布开放权重模型多久?
I would say for a few years.
我会说几年吧。
I think that like in the US, there's not a clear business model for it.
我认为像在美国,它没有清晰的商业模式。
I have been writing about open models for a while and these Chinese companies have realized it.
我写关于开放模型的文章有一段时间了,这些中国公司已经意识到了这一点。
So I get inbound from some of them.
所以我收到他们中一些人的联系。
And they're smart and realize the same constraints, which is that a lot of US tech companies and other IT companies won't pay for an API subscription to Chinese companies for security concerns.
他们很聪明,也意识到同样的限制,那就是很多美国科技公司和其他IT公司出于安全考虑不会向中国公司支付API订阅费用。
This has been a longstanding habit in tech.
这在科技行业一直是个长期习惯。
And the people at these companies then see open weight models as an ability to influence and take part of a huge growing AI expenditure market in the US.
这些公司的人随后将开放权重模型视为一种影响和参与美国巨大增长的AI支出市场的能力。
And they're very realistic about this.
他们对此非常现实。
And it's working for them.
而且这对他们有效。
it's working for them 常用搭配
这对他们有效
表示某个方法或策略对他们产生了预期的好效果。
And I think that the government will see that that is building a lot of influence internationally in terms of uptake of the technology.
而且我认为政府会看到,这在技术采用方面正在国际上建立很大的影响力。
So there's going to be a lot of incentives to keep it going.
所以会有很多激励措施让它继续下去。
But building these models and doing the research is very expensive.
但构建这些模型和做研究非常昂贵。
So at some point I expect consolidation,
所以在某个时候我预计会整合,
but I don't expect that to be a story of 2026,
但我不认为那会是2026年的故事,
where there will be more open model builders throughout 2026 than there were in 2025.
到那时2026年全年的开源模型构建者会比2025年更多。
And a lot of the notable ones will be in China.
而且很多值得注意的会来自中国。
You were going to say something?
你刚才想说什么?
Yes.
是的。
You mentioned DeepSeq losing its crown.
你提到DeepSeq失去了它的王冠。
I do think to some extent, yes.
我确实认为在某种程度上是这样。
to some extent 常用搭配
在某种程度上
用于表示部分同意或部分正确,语气有所保留。
But we also have to consider, though, they are still, I would say, slightly ahead.
但我们也必须考虑到,他们仍然,我会说,稍微领先。
And the other ones, it's not that DeepSeq got worse.
而其他的,并不是DeepSeq变差了。
It's just like the other ones are using the ideas from ZipSeq.
只是其他模型在借鉴ZipSeq的想法。
For example, you mentioned Kimi, same architecture.
例如,你提到的Kimi,相同的架构。
They're training it.
他们正在训练它。
And then again, we have this leapfrogging where they might be at some point in time a bit better because they have the more recent model.
然后我们又有这种跳跃式发展,他们可能在某个时间点稍微更好,因为他们有更新的模型。
And I think this comes back to the fact that there won't be a clear winner.
我认为这又回到了一个事实:不会有明确的赢家。
comes back to the fact that 句型
归根结底是因为……这个事实
come back to the fact that [clause]
用于把讨论引回核心原因或关键事实。
It will just be like that.
就会像那样。
One person releases something, the other one comes in.
一个人发布一些东西,另一个人就来了。
And the most recent model is probably always the best model.
而最新的模型可能总是最好的模型。
Yeah.
是的。
We'll also see that Chinese companies have different incentives.
我们还会看到中国公司有不同的激励措施。
So like DeepSeq is very secretive,
所以像DeepSeq非常神秘,
where some of these startups are like the Minimaxes and Z.AIs of the world.
这些初创公司就像世界上的MiniMax和Z.AI一样。
Those two literally have filed IPO paperwork and they're trying to get Western mindshare and do a lot of outreach there.
那两家确实已经提交了IPO文件,他们正试图在西方赢得关注度,并在那里做大量推广。
So I don't know if these incentives will kind of change the model development
所以我不知道这些激励措施是否会改变模型开发,
because DeepSeq famously is built by a hedge fund, High Flyer Capital.
因为DeepSeq众所周知是由一家对冲基金High Flyer Capital建立的。
And we don't know exactly what we don't know what they use the models for or if they care about this.
我们并不确切知道我们不知道他们用这些模型做什么,或者他们是否在意这个。
They're secretive in terms of communication.
他们在沟通方面很神秘。
They're not secretive in terms of the technical reports that describe how their models work.
但在描述其模型工作原理的技术报告方面,他们并不神秘。
They're still open on that front.
他们在那一方面仍然开放。
on that front 常用搭配
在那一方面
用于指代前面提到的某个方面或领域,表示在该方面的情况。
And we should also say on the Opus 4.5 hype,
我们还应该说说关于Opus 4.5的炒作,
there's the layer of something being the darling of the X echo chamber on Twitter echo chamber
有一层是某样东西成为X上回音室或Twitter回音室的宠儿,
echo chamber 常用搭配
回音室(指信息或观点在封闭圈子里不断重复强化)
常用来形容社交媒体或特定群体中,相同观点被反复传播、缺乏外部声音的现象。
and the actual amount of people that are using the model.
以及实际使用该模型的人数。
I think it's probably fair to say that JGPT and Gemini are focused on the broad user base
我认为可以说JGPT和Gemini专注于广泛的用户群,
it's probably fair to say 句型
可以说,……大概是合理的说法
it's probably fair to say that [clause]
用于表达一个有保留但合理的判断,语气比直接断言更委婉。
that just want to solve problems in their daily lives.
他们只想解决日常生活中的问题。
And that user base is gigantic.
而这个用户群是巨大的。
So the hype about the coding may not be represented by the actual use.
所以关于编程的炒作可能并不代表实际使用情况。
I would say also a lot of the usage patterns are, like you said, name recognition, brand and stuff,
我想说,很多使用模式,就像你说的,是名称认知、品牌之类的,
name recognition 常用搭配
品牌知名度,名字被认出的程度
用于商业或营销语境,指消费者因为听说过某个名字而选择它。
but also muscle memory almost, where, you know, like JGPD has been around for a long time.
但几乎也是肌肉记忆,你知道,就像JGPD已经存在很长时间了。
muscle memory 常用搭配
肌肉记忆(指因长期重复而形成的下意识习惯)
可引申用于形容人们对某事物因长期使用而形成的惯性依赖。
People just got used to using it and it's kind of like almost like a flywheel.
人们只是习惯了使用它,它几乎就像一个飞轮。
got used to 常用搭配
习惯了
表示逐渐适应某事物,常用于描述从陌生到习以为常的过程。
they recommend it to other users and that stuff.
他们把它推荐给其他用户之类的。
One interesting point is also the customization of LLMs.
另一个有趣的点是LLM的定制化。
For example, ChatGPT has a memory feature, right?
例如,ChatGPT有一个记忆功能,对吧?
And so you may have a subscription and you use it for personal stuff,
所以你可能有一个订阅,你用它来处理个人事务,
but I don't know if you want to use that same thing at work, you know, because it's a boundary between private and work.
但我不知道你是否想在工作中使用同样的东西,你知道,因为这是私人和工作之间的界限。
If you're working at a company, they might not allow that or you may not want that.
如果你在一家公司工作,他们可能不允许那样,或者你可能不想要那样。
And I think that's also an interesting point where you might have multiple subscriptions.
我认为这也是一个有趣的点,你可能拥有多个订阅。
One is just clean code. It has nothing of your personal images or hobby projects in there.
一个只是干净的代码。里面没有你的个人图像或爱好项目。
It's just like the work thing.
它就像工作用的东西。
And then the other one is your personal thing.
然后另一个是你的个人东西。
So I think that's also something where two different use cases, and it doesn't mean you only have to have one.
所以我认为这也是两个不同用例的情况,这并不意味着你只能有一个。
I think the future is also multiple ones.
我认为未来也是多个。
What model do you think won 2025?
你认为哪个模型赢得了2025年?
And what model do you think is going to win 26?
你认为哪个模型将赢得26年?
I think in the context of a consumer chatbots is a question of,
我认为在消费者聊天机器人的背景下,这是一个问题,
are you willing to bet on Gemini over ChatGPT?
你愿意押注Gemini而不是ChatGPT吗?
willing to bet on 常用搭配
愿意押注于……,看好……
用于表达对某人或某事物有信心,愿意投入资源或承担风险支持它。
which I would say in my gut feels like a bit of a risky bet
我会说,我的直觉感觉这有点冒险
in my gut 常用搭配
凭直觉,内心深处觉得
用于表达基于直觉而非理性分析的判断,语气较口语化。
because open AI has been the incumbent and there's so many benefits to that in tech.
因为OpenAI一直是现任者,在科技领域这有很多好处。
the incumbent 常用搭配
现任者,现有主导者
商业或政治语境中指当前占据主导地位的一方,常带有需要被挑战的意味。
I think that momentum if you look at 2025 was on gemini side
我认为如果你看2025年,势头是在Gemini这边
but they were starting from such a low point
但他们是从如此低的起点开始的
i think on rip bard and these earlier attempts of getting started
我认为在RIP Bard和这些早期的尝试中
i think huge credit for them for powering through the organizational chaos to make that happen
我认为他们非常值得称赞,能够克服组织混乱实现这一目标
powering through 常用搭配
咬牙坚持挺过,强行推进
表示在困难或混乱中仍然坚持完成某事,强调克服阻力的毅力。
but also it's hard to bet against chat to open ai
但同样很难押注ChatGPT会输给OpenAI
hard to bet against 句型
很难看衰……,很难赌它输
it's hard to bet against [someone/something]
用于表达某人或某事物实力强劲,不看好它是不明智的。
because they always come off cast as so chaotic
因为他们总是显得如此混乱
but they're very good at landing things
但他们非常擅长把事情做成
good at landing things 常用搭配
擅长把事情做成、落地
用于形容一个人或组织能够成功实现目标或完成项目,强调执行力。
And I think like personally, I have very mixed reviews of GPT-5,
而且我认为就个人而言,我对GPT-5的评价非常复杂,
mixed reviews 常用搭配
褒贬不一的评价
原指对产品、电影等的评价有好有坏,可引申为对某事物的复杂看法。
but it had to have saved them so much money with the headline feature being a router where most users are no longer charging their GPU costs as much.
但它一定为他们节省了很多钱,主要功能是一个路由器,大多数用户不再像以前那样承担那么多GPU成本。
So I think it's very hard to dissociate the things that I like out of models versus the things that are going to actually be a general public differentiator.
所以我认为很难将我喜欢模型的地方与真正会成为大众差异化因素的地方分开。
What do you think about 2026? Who's going to win?
你觉得2026年怎么样?谁会赢?
I'll say something, even though it's risky.
我要说一些话,尽管这有风险。
I will say that I think Gemini will continue to take progress on ChatGPT.
我想说,我认为Gemini会在ChatGPT的基础上继续取得进展。
I think Google scale when both of these are operating at such extreme scales.
我认为当这两者都以如此极端的规模运作时,谷歌的规模优势就体现出来了。
And Google has the ability to separate that research and product a bit better.
而谷歌有能力把研究和产品稍微更好地分开。
We hear so much about OpenAI being chaotic operationally and chasing the high impact thing, which is a very startup culture.
我们听到很多关于OpenAI在运营上混乱、追逐高影响力的事情,这是一种非常创业公司的文化。
chasing the high impact thing 常用搭配
追逐高影响力的事情
用于描述组织或个人倾向于追求能产生重大影响的项目,有时暗含不够稳健之意。
And then on the software and enterprise side, I think Anthropic will have continued to success as they've again and again been set up for that.
然后在软件和企业方面,我认为Anthropic会继续取得成功,因为他们一次又一次地为此做好了准备。
And obviously Google's cloud has a lot of offerings, but I think this kind of like Gemini name brand is important for them to build.
显然谷歌云有很多产品,但我认为这种Gemini品牌对他们来说很重要,需要去建立。
And Google's cloud will continue to do well, but that's kind of a more complex thing to explain in the ecosystem
谷歌云会继续表现良好,但这是在生态系统中更复杂的事情,需要解释
because that's competing with the likes of Azure and AWS rather than on the model provider side.
因为那是与Azure和AWS等竞争,而不是在模型提供商方面。
the likes of 常用搭配
像……这样的(人或事物)
用于列举同类中的典型代表,常指知名或重要的人或事物。
Well, infrastructure, you think GPUs give an advantage?
那么,基础设施方面,你认为GPU能带来优势吗?
Largely because the margin on NVIDIA chips is insane.
主要是因为NVIDIA芯片的利润率非常高。
And Google can develop everything from top to bottom to fit their stack and not have to pay this margin.
而谷歌可以从上到下开发一切来适配他们的技术栈,而不必支付这个利润。
And they've had a head start in building data centers.
而且他们在建设数据中心方面有先发优势。
head start 常用搭配
先发优势,领先起步
指在竞争开始前就占据的有利位置,常用于商业、科技等语境。
So all of these things that have both high lead times and very hard margins on high costs,
所以所有这些既交货周期长、又要在高成本上承受极高利润压力的东西,
Google has just kind of a historical advantage there.
谷歌在这方面恰好拥有一种历史性的优势。
And if there's going to be a new paradigm, it's most likely to come from OpenAI,
而如果会出现一种新范式,它最有可能来自 OpenAI,
where their research division again and again has kind of shown this ability to land a new research idea or a product.
因为他们的研究部门一次又一次地展现出这种能力:能够落地一个新的研究想法或产品。
I think deep research, Sora, O1 thinking models, all these definitional things have come from OpenAI,
我认为深度研究、Sora、O1 思维模型,所有这些具有定义意义的东西都来自 OpenAI,
and that's got to be one of their top traits as an organization.
而这一定是他们作为一个组织的最大特质之一。
So it's kind of hard to bet against that,
所以很难去赌他们不行,
but I think a lot of this year will be about scale and optimizing what could be described as low-hanging fruit in models.
但我认为今年很大程度上将围绕规模扩展,以及优化模型中可被称为“低垂果实”的部分。
low-hanging fruit 常用搭配
低垂的果实(指容易实现的目标或容易取得的成果)
比喻最容易达成或最显而易见的机会,常用于商业和策略讨论。
And clearly there's a trade-off between intelligence and speed.
而且显然,智能和速度之间存在权衡。
trade-off 常用搭配
权衡,取舍
指在两种不可兼得的事物之间做出的取舍,常用于描述利弊权衡。
This is what ChatGPT 5 was trying to solve behind the scenes.
这就是 ChatGPT 5 在幕后试图解决的问题。
behind the scenes 常用搭配
在幕后,不为人知地
用于描述不为公众所见的工作或过程,强调隐蔽性。
It's like, do people actually want intelligence, the broad public, or do they want speed?
就像是在问,大众真的想要智能,还是想要速度?
I think it's a nice variety, actually, or the option to have a toggle there.
其实我觉得提供多样性,或者说有一个可以切换的选项,是件好事。
I mean, first, for my personal usage, most of the time when I look something up, I use ChatGPT to ask a quick question, get the information, I want it fast.
我的意思是,首先,就我个人使用而言,大多数时候我查东西时,会用 ChatGPT 问个快速问题、获取信息,我希望它快。
look something up 常用搭配
查阅,查找(信息)
指通过书籍、网络等查找所需信息,常用于日常口语。
For, you know, most daily tasks, I use the quick model.
对于大多数日常任务,我会用快速模型。
Nowadays, I think the auto mode is pretty good where you don't have to specifically say thinking or, you know, non-thinking and stuff.
如今,我觉得自动模式相当不错,你不需要特意说“思考”或者“非思考”之类的。
Then again, I also sometimes want the pro mode.
不过话说回来,我有时也想要专业模式。
Very often what I do is when I have something written, I put it into a chatubity and say, hey, do a very thorough check.
很多时候我会这样做:当我写好东西后,我把它放进一个聊天工具里,说,嘿,做一个非常彻底的检查。
Are all my references correct? Are all my thoughts correct? Did I make any formatting mistakes? And are the figure numbers wrong or something like that?
我所有的引用都正确吗?我所有的想法都正确吗?我有没有犯任何格式错误?还有图表编号是不是错了之类的?
And I don't need that right away. It's something, okay, I finish my stuff, maybe have dinner, let it run, come back, and it goes through this.
而且我并不需要马上得到结果。就是那种,好吧,我做完我的事,也许去吃个晚饭,让它跑着,回来之后,它就把这些过一遍。
And I think, see, this is where I think it's important to have this option.
我觉得,你看,这就是我认为拥有这个选项很重要的地方。
I would go crazy if for each query I would have to wait 30 minutes or 10 minutes.
如果每次查询我都得等30分钟或10分钟,我会疯掉的。
go crazy 地道口语
发疯,受不了
口语中夸张地表达因等待、压力等而极度不耐烦或崩溃。
That's me. I'm like sitting over here, losing my mind, that you use the router and the non-thinking model.
这就是我。我就像坐在这儿,快疯了,你居然用路由器和非思考模型。
losing my mind 地道口语
快疯了,抓狂
口语中夸张表达极度焦虑、沮丧或无法忍受的状态。
I'm like, how do you live with that? It's like my reaction.
我就像在说,你怎么受得了?这就是我的反应。
how do you live with that 句型
你怎么受得了?
how do you live with [something]
用于表达对他人做法或处境感到难以置信或无法接受,带有调侃或夸张语气。
I've been heavily on Chat to BT for a while.
我有一阵子一直在重度使用Chat to BT。
Never touched five non-thinking.
从没碰过5的非思考模式。
I find its tone and then its propensity of errors.
我发现它的语气,还有它容易出错的倾向。
It's just like it has a higher likelihood of errors.
就好像它出错的概率更高。
Some of this is from back when opening, I released 03, which was the first model to do this deep search and find many sources and integrate them for you.
其中一部分来自当初发布的时候,我发布了03,那是第一个能做这种深度搜索、找到很多来源并为你整合的模型。
So I became habituated with that.
所以我就习惯了那样。
became habituated with 常用搭配
逐渐习惯、适应了某事
用于描述经过一段时间后对某种做法或状态习以为常,较正式。
So I will only use GPT 5.2 thinking or pro
所以我只会用 GPT 5.2 的思考模式或专业模式
when I'm finding any sort of information query for work,
当我查找任何与工作相关的信息查询时,
whether that's a paper or some code reference that I found.
无论是论文还是我找到的某些代码参考。
And it's just like, I will regularly have like five pro queries going simultaneously,
就像这样,我通常会同时运行五个专业查询,
each looking for one specific paper or feedback on an equation or something.
每个查询都在寻找一篇特定的论文,或者关于某个方程的反馈之类的。
I have a fun example where I just needed to answer as fast as possible.
我有一个有趣的例子,当时我需要尽快给出答案。
For this podcast, before I was going on the trip,
为了这个播客,在我出发旅行之前,
I have a local GPU running at home and I wanted to run a long RL experiment.
我家里有一台本地 GPU 在运行,我想跑一个长时间的强化学习实验。
And usually I also unplug things because
通常我也会拔掉一些东西,因为
you never know if you're not at home, you don't want to have things plugged in.
你永远不知道,如果你不在家,你不想让东西一直插着电。
you never know 地道口语
你永远说不准、无法预料
口语中用来表示某事无法确定,常引出谨慎行事的理由。
And I accidentally unplugged the GPU.
我不小心拔掉了 GPU 的电源。
It was like my wife was already in the car and it's like oh dang
就像我妻子已经在车里了,然后就像,哦糟糕
and then basically i wanted as fast as possible a bash script that runs my different uh experiments and evaluation and
然后基本上我想要尽快写一个 bash 脚本,用来运行我不同的呃实验和评估,然后
i did something i know i learned how to use the bash uh interface or bash terminal
我做了一些我知道的事情,我学会了如何使用 bash 呃界面或 bash 终端
but in that moment i just needed like 10 seconds
但在那一刻,我只需要大约 10 秒钟
give me the command this is a hilarious situation but yes
给我命令,这情况太搞笑了,但没错
what did you use so i did the non-thinking fastest model it gave me
你用了什么,所以我用了非思考模式的最快模型,它给了我
The bash command I to chain different scripts to each other and then
那个 bash 命令,我把不同的脚本串联起来,然后
the thing is like you have the t thing where you want to route this to a log file top of my head
问题是,就像你有那个 t 命令,你想把这个路由到一个日志文件,我脑子里首先想到的
top of my head 常用搭配
(凭记忆)立刻想到的
常用于 off the top of my head 结构,表示未经查证、凭印象说出。
I was just like in a hurry I could have thought about it myself by the way
我当时就是很匆忙,顺便说一句,我本可以自己想想的
in a hurry 常用搭配
匆忙、赶时间
描述因时间紧迫而仓促行事的状态。
I don't know if there's a representative case why you wait in the car you have to run you unplug the GPU you have to generate a bash script
我不知道是否有代表性案例,为什么你要在车里等,你得跑,你拔掉 GPU,你得生成一个 bash 脚本
This sounds like a movie like in which possible I use Gemini for that
这听起来像一部电影,就像可能我用 Gemini 来做那个
So I use thinking for all the information stuff and then Gemini for fast things or stuff that I could sometimes Google
所以对于所有信息类的东西,我用 thinking,然后对于快速的事情,或者我有时可以谷歌的东西,用 Gemini
which is like it's good at explaining things and I trust that it has this kind of background of knowledge and it's simple and the Gemini app has got a lot better and it's good for that sort of things
它就像,它擅长解释事情,我相信它有这种知识背景,而且它很简单,Gemini 应用已经好多了,它适合那种事情
and then for code and any sort of philosophical discussion I use Claude Opus 4.5 also always with extended thinking extended thinking and inference time scaling
然后对于代码和任何类型的哲学讨论,我用 Claude Opus 4.5,也总是用扩展思考、扩展思考和推理时间缩放
is just a way to make the models marginally smarter and I will always edge on that side when the progress is very high
只是一种让模型稍微更聪明的方法,当进展非常快时,我总是会偏向那一边
because you don't know when that'll unlock a new use case
因为你不知道什么时候那会解锁一个新的使用场景
and then sometimes use grok for real-time information
然后有时候用Grok获取实时信息
or finding something on AI Twitter
或者在AI推特上找东西
that I knew I saw and I need to dig up and I just fixated on
我知道我见过,我需要挖出来,而我只是盯着
although when Grok 4 came out, the Grok 4, what is super heavy
尽管当Grok 4出来时,Grok 4,什么是超重型
which was like their pro variant was actually very good
那就像是他们的专业版,实际上非常好
and I was pretty impressed with it
我对它印象相当深刻
and I just kind of like muscle memory lost track of it
而我只是有点肌肉记忆,跟丢了它
muscle memory 常用搭配
肌肉记忆,指因反复练习而形成的下意识习惯
比喻长期形成的习惯性动作或选择,无需刻意思考。
with having the chat to BT app open
因为一直开着聊天到BT应用
so I use many different things
所以我用很多不同的东西
yeah I actually do use Grok 4 heavy for debugging
是的,我实际上确实用Grok 4重型来调试
for like hardcore debugging
用于像硬核调试
and the other ones can't solve it I find that it's the best
而其他的解决不了,我发现它是最好的
and I it's interesting because you say Chad GPT is the best interface for me for that same reason
而且我,有趣的是,你说Chad GPT对我来说是最好的界面,出于同样的原因
but this could be just momentum
但这可能只是惯性
Gemini is the better interface for me I think
我认为Gemini对我来说是更好的界面
because I fell in love with their best needle in the haystack
因为我爱上了他们最好的大海捞针
fell in love with 常用搭配
爱上了、非常喜欢上
口语中常用来表示对某物或某功能产生强烈好感。
if I ever put something that has a lot of context
如果我曾经放入一些有很多上下文的东西
but I'm looking for very specific kinds of information
但我正在寻找非常具体的信息
make sure it tracks all of it
确保它追踪所有内容
I find at least the Gemini for me has been the best.
我发现至少Gemini对我来说一直是最好的。
So it's funny with some of these models, if they win your heart over for one particular feature on a one particular day for that particular query, that prompt,
所以有趣的是,对于其中一些模型,如果它们在某一天、针对某个特定的查询、那个提示词,用某个特定功能赢得了你的心,
win your heart over 常用搭配
赢得你的心、让你喜欢上
用于描述某事物通过某个优点打动了你,使你产生偏爱。
you're like, this model is better.
你就会觉得,这个模型更好。
And so you'll just stick with it for a bit until it does something really dumb.
于是你就会继续用它一阵子,直到它做出一些非常愚蠢的事。
stick with it 常用搭配
继续坚持用它、不换别的
口语中表示在尝试后决定继续使用某个选择。
There's like a threshold effect, some smart thing, and then you fall in love with it.
就像有一种阈值效应,它做了些聪明的事,然后你就爱上它了。
And then it does some dumb thing.
然后它又做了些愚蠢的事。
And you're like, you know what? I'm going to switch and try clawed and JGBT and all that kind of stuff.
然后你会说,你知道吗?我要换一个,试试Clawed和JGBT之类的。
you know what 地道口语
你知道吗、我跟你说
口语中用于引出决定或坦白,起过渡和强调作用。
This is exactly like you use it until it breaks, until you have a problem, and then you change the LM.
这就像你用一样东西,直到它坏掉,直到你遇到问题,然后你就换掉那个语言模型。
And I think it's the same how we use anything, like our favorite text editor, operating systems, or the browser.
我觉得我们使用任何东西都是这样,比如我们最喜欢的文本编辑器、操作系统,或者浏览器。
I mean, there are so many browser options, Safari, Firefox, Chrome, all the, kind of relatively similar,
我的意思是,浏览器选项太多了,Safari、Firefox、Chrome,都差不多,
but then there are edge cases, maybe extensions you want to use, and then you switch.
但然后会有一些边缘情况,也许是你想用的扩展,然后你就换了。
But I don't think there is any one who types the same thing, like the website, into different browsers and compares them.
但我觉得没有人会把同样的东西,比如网站,输入到不同的浏览器里然后比较它们。
You only do that when the website doesn't render, if something breaks, I think.
我觉得你只有在网站无法渲染、出问题的时候才会那么做。
So that's a good point. I think you use it until it breaks, and then you explore other options, I think.
所以这是个好观点。我觉得你用一样东西直到它坏掉,然后你再探索其他选项,我觉得是这样。
On the long context thing, I was also a Gemini user for this.
关于长上下文这件事,我在这方面也是Gemini的用户。
But the GPT-5.2 release blog had crazy long context scores
但GPT-5.2的发布博客有疯狂的长上下文分数
where a lot of people were like, did they just figure out some algorithmic change?
很多人都在想,他们是不是只是搞出了某种算法上的改变?
It went from 30% to 70% or something in this minor model update.
在这个小版本模型更新中,它从30%左右涨到了70%左右。
So it's also very hard to keep track of all of these things.
所以也很难跟踪所有这些事情。
keep track of 常用搭配
跟踪、掌握……的动态
用于表示持续关注并记录某事物的变化或进展。
But now I look more favorably at GPT-5.2's long context.
但现在我更看好GPT-5.2的长上下文。
So it's just kind of like, how do I actually get to testing this never-ending battle?
所以这就像,我到底要怎么去测试这场永无止境的战斗?
It's interesting that none of us talked about the Chinese models from a user usage perspective.
有趣的是,我们没有人从用户使用角度谈论中国模型。
What does that say? Does that mean the Chinese models are not as good?
这说明什么?是不是意味着中国模型没那么好?
Or does that mean we're just very biased and U.S.-focused?
还是说我们只是非常偏颇、只关注美国?
I do think that that's currently the discrepancy between just the model and the platform.
我确实认为,这就是目前模型和平台之间的差异。
So I think the open models, they are more known for the open weights, not the platform yet.
所以我认为,开源模型更出名的是开放权重,而不是平台。
There are also a lot of companies that are willing to sell you the open model inference at a very low cost.
也有很多公司愿意以非常低的成本向你出售开源模型的推理服务。
I think like Open Router, it's easy to do the look at multi-model things.
我觉得像Open Router,很容易查看多模型的东西。
You can run DeepSeq on perplexity.
你可以在Perplexity上运行DeepSeq。
I think all of us sitting here are like, we use OpenAI GPT-5 Pro consistently.
我觉得我们坐在这里的所有人都一样,我们一直使用OpenAI GPT-5 Pro。
We're all willing to pay for the marginal intelligence gain.
我们都愿意为那一点边际智能提升付费。
And anyone that's like these models from the U.S. are better.
而任何认为这些来自美国的模型更好的人。
And in terms of the outputs, I think that the question is, will they stay better for this year and for years going?
就输出而言,我认为问题是,它们在今年以及未来几年会保持更好吗?
in terms of 常用搭配
就……而言;在……方面
用于引出讨论的某个方面或角度,正式和口语中都很常见。
for years going 常用搭配
未来几年;往后几年
口语中表示从今往后的若干年,相当于 for years to come。
But it's like so long as they're better, I'm going to pay for it to use them.
但就像只要它们更好,我就会付费使用它们。
so long as 常用搭配
只要
引导条件从句,表示只要满足某条件,主句就成立,比 as long as 稍正式。
I think there's also analysis that shows that like the.
我认为也有分析表明,就像那个。
way that the chinese models are served this
中国模型被服务的方式,这
you could argue due to expert controls or not is that they use fewer gpus for replica which makes them slower and have different errors and it's like speed and intelligence if these things are in your favor as a user i think in the u.s a lot of users will go for this and i think that that is something that will spur these chinese companies to want to compete in other ways whether it's like free or substantially lower costs or it'll breed creativity in terms of offerings, which is good for the ecosystem.
你可以争辩说,由于专家控制与否,它们为副本使用更少的GPU,这使它们更慢,并且有不同的错误,这就像速度和智能,如果这些对你作为用户有利,我认为在美国很多用户会倾向于这个,我认为这会促使这些中国公司想以其他方式竞争,无论是免费还是大幅降低成本,或者会在产品方面激发创造力,这对生态系统有好处。
you could argue 句型
你可以说;可以认为
you could argue [that clause]
用于提出一种可讨论、未必确定的观点,语气委婉。
in your favor 常用搭配
对你有利
表示某条件或情况对某人有利,常用于讨论利弊。
But I just think the simple thing is the U.S. models are currently better
但我只是认为简单的事实是,美国模型目前更好
and we use them and I try these other open models and I'm like, fun, but I don't go back to it.
我们使用它们,我尝试这些其他开源模型,我觉得,有趣,但我不会回去用它。
go back to 常用搭配
回去再用;重新使用
表示放弃新尝试后重新回到原来使用的工具或做法。
We didn't really mention programming. That's another use case that a lot of people deeply care about.
我们真的没有提到编程。那是另一个很多人深切关心的用例。
use case 常用搭配
用例;应用场景
指某项技术或产品被实际使用的具体场景,科技讨论中常用。
care about 常用搭配
关心;在意
表示对某事重视或感兴趣,deeply care about 强调非常在意。
So I use basically half and half cursor and clog code
所以我基本上是一半一半地用 Cursor 和 Clog Code
half and half 常用搭配
各一半;一半一半
表示两种事物各占一半比例,口语中很常用。
because I find them to be fundamentally different experience and both useful.
因为我发现它们本质上是不同的体验,而且都很有用。
What do you guys, you program quite a bit.
你们呢,你们编程挺多的。
So what do you use?
那你们用什么?
What's the current vibe?
现在流行什么?
What's the current vibe? 地道口语
现在是什么风潮/氛围?
口语中询问当前流行什么或大家现在的感受,vibe 指氛围或潮流。
So I use the Codex plugin for VS Code.
所以我用 VS Code 的 Codex 插件。
You know, it's very convenient. It's just like a plugin.
你知道,它非常方便。就像一个插件。
And then it's a chat interface that has access to your repository.
然后它是一个可以访问你代码库的聊天界面。
I know that Cloud Code is, I think, a bit different.
我知道 Cloud Code 我觉得有点不同。
It's a bit more agentic. It touches more things. It does a whole project for you.
它更自主一些。它涉及更多东西。它为你完成整个项目。
a bit more 常用搭配
稍微更……一些
用于缓和语气,表示程度略高,口语中常用来弱化比较。
I'm not quite there yet where I'm comfortable with that because maybe I'm a control freak,
我还没到能接受那种方式的程度,因为也许我是个控制狂,
I'm not quite there yet 地道口语
我还没到那个程度
表示自己尚未准备好接受或做到某事,语气委婉。
control freak 地道口语
控制狂
口语中自嘲或形容喜欢掌控一切细节的人。
but I still would like to see a bit what's going on.
但我还是想稍微看看发生了什么。
Codex is kind of like right now for me, like the sweet spot where it is helping me,
Codex 现在对我来说有点像最佳平衡点,它在帮助我,
sweet spot 地道口语
最佳平衡点;理想状态
指各方面恰到好处、最理想的位置或状态。
but it is not taking completely over.
但它并没有完全接管。
I should mention one of the reasons I do use Cloud Code is to build the skill of programming with English.
我应该提一下,我使用 Cloud Code 的原因之一是为了培养用英语编程的技能。
I mean, the experience is fundamentally different.
我的意思是,体验是根本不同的。
You're, as opposed to micromanaging the details of the process of the generation of the code
你是,而不是去微观管理代码生成过程的细节
as opposed to 常用搭配
而不是;与……相对
用于对比两种不同做法或情况,强调区别。
and looking at the diff, which you can in cursor, if that's the idea you use,
然后查看差异,如果你用的是 Cursor 的话,你可以在里面看差异,
and in changing, altering, looking and reading the code and understanding the code deeply as you progress
并且在修改、改动、查看和阅读代码的过程中,随着进展深入理解代码
versus just kind of like thinking in this design space and just guiding it at this macro level,
而不是只是在这个设计空间里思考,只是在这个宏观层面上引导它,
which I think is another way of thinking about the programming process.
我认为这是思考编程过程的另一种方式。
Also, we should say that Cloud Code, it just seems to be somehow a better utilization of Cloud Opus 4.5.
另外,我们应该说,Cloud Code 似乎以某种方式更好地利用了 Cloud Opus 4.5。
It's a good side-by-side for people to do.
这是一个很适合人们并排对比使用的工具。
side-by-side 常用搭配
并排对比
指把两个事物放在一起比较,常用于工具或产品对比。
So you can have cloud code open, you can have cursor open, you can have VS code open, and you can select the same models on all of them and ask questions.
所以你可以打开 Cloud Code,可以打开 Cursor,可以打开 VS Code,然后在它们上面选择相同的模型并提问。
It's very interesting.
这非常有意思。
Like the cloud code is way better in that domain.
就像 Cloud Code 在那个领域要好得多。
way better 地道口语
好得多
口语中 way 用来加强比较级,表示程度明显更高。
It's remarkable.
这很了不起。
All right. We should say that both of you are legit on multiple fronts.
好的。我们应该说,你们两位在多个方面都很厉害。
on multiple fronts 常用搭配
在多个方面
表示在多个领域或方面都具备某种特质,front 指方面或战线。
Researchers, programmers, educators, tweeterers, and on the book front, too.
研究人员、程序员、教育者、推特用户,在写书方面也是如此。
So Nathan, at some point soon, hopefully has an RLHF book coming out.
所以 Nathan 希望不久之后能出一本关于 RLHF 的书。
at some point 常用搭配
在某个时候
表示不确定具体时间但将来会发生的某个时刻。
It's available for pre-order and there's a full digital pre-print
它可以预订,还有完整的数字预印本
available for pre-order 常用搭配
可以预订
用于商品或书籍尚未正式发售但已可提前订购。
just making it pretty and better organized for the physical thing,
只是为了让实体书更精美、更有条理,
which is a lot of why I do it because it's fun to create things that you think are excellent in the physical form
这也是我这么做的很大原因,因为创造你认为在实体形式上很出色的东西很有趣
when so much of our life is digital.
当我们的生活如此数字化时。
I should say, going to perplexity here, Sebastian Roshka is a machine learning researcher and author known for several influential books.
我应该提一下,来到Perplexity这里,Sebastian Roshka是一位机器学习研究员和作家,以几本有影响力的书闻名。
known for 常用搭配
以……闻名
表示某人或某物因某种特点或成就而出名。
A couple of them that I wanted to mention, which is a book I highly recommend, Build a large language model from scratch and the new one build a reasoning model from scratch
我想提几本,其中一本我强烈推荐,《从零构建大型语言模型》,还有新书《从零构建推理模型》
from scratch 常用搭配
从零开始;从头做起
表示不借助现成基础,完全从头构建或制作。
so i'm really excited about that building stuff from scratch
所以我对从零构建东西感到非常兴奋
is one of the most powerful ways of learning honestly
老实说,这是最强大的学习方式之一
building an element from scratch is a lot of fun it's also a lot of to learn and like you said it's probably the best way to learn how something really works
从零构建一个元素很有趣,也能学到很多,就像你说的,这可能是了解某事物真正运作方式的最佳途径
because you can look at figures but figures can have mistakes
因为你可以看图表,但图表可能有错误
you can look at concepts explanations but you might misunderstand them.
你可以看概念解释,但你可能误解它们。
But if you see there is code and the code works, you know it's correct.
但如果你看到有代码,而且代码能运行,你就知道它是正确的。
I mean, there's no misunderstanding. It's like, it's precise. Otherwise it wouldn't work.
我的意思是,没有误解。就像,它是精确的。否则它就不会运行。
And I think that's like kind of like the beauty behind coding.
我觉得这就像是编程背后的美妙之处。
It is kind of like, it doesn't lie.
它有点像是,它不会说谎。
It's math basically.
基本上就是数学。
So even though with math, I think you can have mistakes in a book you would never notice because you're not running the math when you are reading the book.
所以即使有数学,我觉得书里可能有你永远不会注意到的错误,因为你读书的时候并没有在运行数学。
You can't verify this.
你没法验证这一点。
And with code, what's nice is you can verify it.
而有了代码,好处就是你可以验证它。
Yeah, I agree with you about the LM from scratch book,
是的,我同意你关于《从零构建语言模型》这本书的看法,
agree with 常用搭配
同意(某人/某观点)
表示赞同某人的看法或说法,后接人或观点。
it's nice to tune out everything else, the internet and so on, and just focus on the book.
能屏蔽掉其他一切,比如互联网之类的,只专注于这本书,这很好。
tune out 常用搭配
屏蔽;不去理会
表示有意忽略周围的干扰或噪音,专注于某事。
focus on 常用搭配
专注于
表示把注意力集中在某事物上。
But, you know, I read several like, you know, history books.
但是,你知道,我读过好几本,你知道,历史书。
It's just less lonely somehow.
不知怎的,就是没那么孤独。
It's really more fun.
真的更有趣。
Like, for example, on the programming front, I think it's genuinely more fun to program with an LLM.
比如,在编程方面,我觉得用大语言模型编程真的更有趣。
And I think it's genuinely more fun to read with an LLM.
而且我觉得用大语言模型阅读真的更有趣。
But you're right, like this distraction should be minimized.
但你说得对,这种分心应该尽量减少。
So you use the LLM to basically enrich the experience, maybe add more context.
所以你用大语言模型来丰富体验,也许增加更多背景信息。
Maybe the rate of aha moments for me in a small scale is really high with LLMs.
也许在小规模上,我获得顿悟时刻的频率真的很高,有了大语言模型。
aha moments 地道口语
顿悟时刻;恍然大悟的瞬间
指突然理解或想通某事的时刻,口语中常用。
100%. I also want to correct myself.
百分之百。我也想纠正一下自己。
correct myself 常用搭配
纠正自己
表示说话人意识到刚才说错,主动更正自己的话。
I'm not suggesting not to use LLMs.
我不是建议不要使用大语言模型。
I suggest doing it in multiple passes, like one pass just offline focus mode.
我建议分多次进行,比如一次只是离线专注模式。
multiple passes 常用搭配
多遍;多次处理
指对同一材料分多次进行阅读或处理,pass 指一遍。
And then after that, I mean, I also take notes,
然后在那之后,我是说,我也会做笔记,
but I try to resist the urge to immediately look things up.
但我尽量克制住立刻去查东西的冲动。
resist the urge 常用搭配
克制冲动
表示忍住想做某事的强烈欲望,后常接 to do。
I do a second pass. It's just like for me more structured this way.
我会做第二遍。对我来说这样只是更有条理。
And I get, I mean, sometimes things are answered in the chapter,
而且我明白,我是说,有时候有些东西在章节里就有答案,
but sometimes also it just helps to let it sink in and think about it.
但有时候,让它慢慢沉淀下来、去思考一下,也很有帮助。
sink in 常用搭配
慢慢被理解;逐渐领会
表示信息或想法需要时间才能被充分消化理解。
Other people have different preferences. I would highly recommend using LLMs when reading books.
其他人有不同的偏好。我非常推荐在读书时使用大语言模型。
highly recommend 常用搭配
强烈推荐
表示非常建议某人做某事或使用某物,语气较强。
For me, it's just, it's not the first thing to do. It's like the second pass.
对我来说,它只是,它不是第一件要做的事。它像是第二遍。
My way of recommendation is to say, I do the opposite.
我推荐的方式是说,我做的正好相反。
I like to use the LLM at the beginning to lay out the full context of like, what is this world that I'm now stepping into?
我喜欢在一开始就用大语言模型来铺开完整的背景,比如,我现在踏入的这个世界是什么?
lay out 常用搭配
铺开;详细说明
表示把信息、背景或计划清楚地呈现出来。
But I try to avoid clicking out of the LLM into the world of like Twitter and blogs.
但我尽量不点出大语言模型,进入像推特和博客那样的世界。
And because then you're now down this rabbit hole, you're reading somebody's opinion.
因为那样你就掉进了这个兔子洞,你在读某个人的观点。
down this rabbit hole 地道口语
陷入一个越查越深的探究之中
形容开始研究某话题后越陷越深、难以抽身,口语常用。
There's a flame war about a particular topic.
有一个关于某个特定话题的激烈骂战。
flame war 常用搭配
网络上的激烈骂战
指网上围绕某话题互相攻击的激烈争论,多含贬义。
And all of a sudden you're no longer, you're now in the realm of the Internet and Reddit and so on.
然后突然之间你就不再是,你现在进入了互联网和Reddit等等的领域。
all of a sudden 常用搭配
突然间
口语中引出意外发生的变化,比 suddenly 更随意。
But if you're purely letting the LLM give you the context of why this matters, what are the big picture ideas?
但如果你纯粹让大语言模型给你提供背景,说明为什么这很重要,大局观是什么?
the big picture 常用搭配
整体情况、大局
指从宏观角度看待问题,而非纠结细节。
But sometimes books themselves are good at doing that, but not always.
但有时候书本身就擅长做这件事,但并非总是如此。
That's why I like the ChatGPT app.
这就是为什么我喜欢ChatGPT应用。
It gives the AI a home in your computer when you can focus on it rather than just being another tab in my mess of Internet options.
它给AI在你的电脑里安了个家,让你能专注于它,而不是只把它当成我那一堆乱七八糟的互联网选项里的又一个标签页。
rather than 句型
而不是
[A] rather than [B]
用于对比两个选项,强调选择前者而非后者。
And I think Cloud Code in particular does a good job of making that a joy, where it seems very engaging as a product.
而且我觉得Cloud Code在这方面做得特别好,让它变得很愉快,作为一个产品它看起来非常吸引人。
does a good job of 常用搭配
在某方面做得很好
评价某人或某产品在某事上表现出色,中性偏褒。
designed to be an interface that your ai will then go out into the world and it's something that is very kind of intangible between it and codex is that it just feels kind of warm and engaging
它被设计成一个界面,让你的AI随后走向世界,而它和Codex之间有种很微妙的东西,就是它感觉起来有点温暖、很吸引人。
where codex can often be as good from open ai but it just kind of like feels a little bit rougher on the edges
而Codex虽然来自OpenAI,往往也一样好,但它就是感觉边缘有点粗糙。
whereas cloud code is makes it fun to build things particularly from scratch
而Cloud Code则让构建东西变得有趣,尤其是从零开始。
from scratch 常用搭配
从零开始
指不借助已有基础,从头做起。
where you just don't like you don't have to care but you trust that it'll make something like obviously good for websites and kind of refreshing tooling and stuff like this which i use it for data analysis
在这里你完全不用操心,但你可以相信它会做出一些显然不错的东西,比如网站和那种让人耳目一新的工具之类的,我就是用它来做数据分析的。
so i my blog we scrape hugging face we keep the download numbers for every data set and model over time now
所以我的博客会抓取Hugging Face的数据,我们持续记录每个数据集和模型随时间变化的下载量。
so we have them and it's like claude was
所以我们有这些数据,就像Claude当时……
Just like, yeah, I've made use of that data.
就像,是的,我利用了那些数据。
made use of 常用搭配
利用了
表示把某资源派上用场,比 use 更强调有目的地利用。
No problem. And I was like, that would have taken me days.
没问题。然后我想,那本来会花我好几天。
And I was like, then I have enough situational awareness to be like, okay, these trends obviously make sense.
然后我想,那我有足够的情境意识去说,好吧,这些趋势显然合理。
make sense 常用搭配
说得通、合理
表示某事逻辑上成立或可以理解,日常高频。
And you can check things because that's just a kind of wonderful interface
而且你可以查看东西,因为那只是一个很棒的界面
where you can have an intermediary and not have to do the kind of awful low-level work
在那里你可以有一个中介,而不必做那种糟糕的低级工作
that you would have to do to maintain different web projects and do this stuff all right
你为了维护不同的网络项目并做这些事情而不得不做的工作,好吧
So we just talked about a bunch of the closed weight models.
所以我们刚刚谈了一堆闭源模型。
Let's talk about the open ones. Uh, so tell me about the landscape of open LM models,
我们来谈谈开源的。呃,那么告诉我开源语言模型的情况,
which are interesting ones, which stand out to you and why. We already mentioned DeepSeek.
哪些是有趣的,哪些对你来说很突出,为什么。我们已经提到了DeepSeek。
stand out 常用搭配
突出、显眼
指在同类中格外引人注意或表现优异。
Do you want to see how many we can name off the top of our head? Yeah, without looking at notes.
你想看看我们能不假思索地说出多少个吗?是的,不看笔记。
off the top of our head 地道口语
不假思索地、凭记忆
表示没有查资料、凭印象说出,口语常用。
DeepSeek, Kimi, MiniMax, Z.ai, Antling, we're just going Chinese. Um, let's throw in Mistral AI, Gemma, um, yeah, GPT-OSS.
DeepSeek、Kimi、MiniMax、Z.ai、Antling,我们只是在说中国的。嗯,让我们加上Mistral AI、Gemma,嗯,是的,GPT-OSS。
The open source model by Jet GPT, actually Nvidia Nemotron had a, Nvidia had a really cool one, Nemotron 3.
Jet GPT的开源模型,实际上Nvidia Nemotron有一个,Nvidia有一个非常酷的,Nemotron 3。
Um, there's a lot of stuff especially at the end of the year.
嗯,有很多东西,尤其是在年底。
Quen? One, maybe the one... Oh yeah, Quen was the obvious name.
Quen?一,也许是那个……哦对,Quen 是那个显而易见的名字。
I was trying to get through that you can get at least 10 Chinese and at least 10 Western.
我当时想表达的是,你至少能列出10个中国的和至少10个西方的。
I think that, I mean, opening... I released their first open model since GPT-2.
我觉得,我是说,开放……我发布了他们自 GPT-2 以来的第一个开放模型。
That was when I, when I meant to talk when I was writing about opening eyes, open model release.
那是在我,当我想谈的时候,我当时在写关于开放视野、开放模型发布的内容。
They're all like, don't forget about GPT-2, which I thought was really funny.
他们都在说,别忘了 GPT-2,我觉得这真的很好笑。
Because it's just such a different time.
因为那完全是一个不同的时代。
But GPT-OSS is actually a very strong model.
但 GPT-OSS 实际上是一个非常强大的模型。
And does some things that the other models don't do very well.
而且它能做一些其他模型做不好的事情。
And I think that, selfishly, I'll promote a bunch of like Western companies.
而且我觉得,出于私心,我会推广一堆像西方公司这样的。
So both in the U.S., in Europe, have these like fully open models.
所以美国和欧洲都有这些完全开放的模型。
So I work at Allen Institute for AI. We've been building ULMO, which releases data and code and all of this.
所以我在艾伦人工智能研究所工作。我们一直在构建 ULMO,它发布数据和代码以及所有这些。
And now we have actual competition for people that are trying to release everything.
而现在,对于那些试图发布一切的人来说,我们有了真正的竞争。
So that other people can train these models.
这样其他人就能训练这些模型。
So there's the Institute for Foundation Models, or /LM 360.
所以有基础模型研究所,或者叫 /LM 360。
Which is like had their K2 models of various types.
它就像是有各种类型的 K2 模型。
APRODIS is a Swiss research consortium.
APRODIS 是一个瑞士研究联盟。
Hugging Face has Small LM, which is very popular.
Hugging Face 有 Small LM,非常受欢迎。
where NVIDIA's Neymatron has started releasing data as well.
英伟达的Neymatron也开始发布数据了。
And then Stanford's Marin Community Project, which is kind of making it so there's a pipeline for people to open a GitHub issue and implement a new idea and then have it run in a stable language modeling stack.
然后是斯坦福的Marin社区项目,它某种程度上是在打造一条流程,让人们可以提交GitHub issue、实现一个新想法,然后让它在稳定的语言建模技术栈上运行。
So this space... That list was way smaller in 2024.
所以这个领域……那份名单在2024年要小得多。
So I think it was like just AI2.
所以我觉得当时基本上就只有AI2。
So that's a great thing for more people to get involved in to understand language models,
所以这是件好事,能让更多人参与进来、理解语言模型,
get involved in 常用搭配
参与、投身于
表示加入某项活动或事业,常与 in 搭配。
which doesn't really have a Chinese company that has an analog.
而这一点其实并没有哪家中国公司能与之对应。
While I'm talking, I'll say that the Chinese open language models tend to be much bigger,
趁我讲的时候,我要说中国的开源语言模型往往要大得多,
and that gives them this higher peak performance as MOEs,
这让它们作为MoE模型能达到更高的峰值性能,
where a lot of these things that we like a lot, whether it was Gemma and Nemetron,
而我们很喜欢的很多东西,无论是Gemma还是Nemetron,
tended to be smaller models from the U.S., which is starting to change from the U.S. and Europe.
往往都是美国的小模型,而这一点正开始从美国和欧洲发生变化。
Mistral Large 3 came out, which was a giant MOE model, very similar to DeepSeek architecture in December.
Mistral Large 3发布了,那是个巨大的MoE模型,架构和12月的DeepSeek非常相似。
And then a startup, RCAI and both Nemetron and NVIDIA have teased MOE models of way bigger than 100 billion parameters, like this 400 billion parameter range coming in this Q1 2026 timeline.
然后还有一家初创公司RCAI,而且Nemetron和英伟达都预告了远超1000亿参数的MoE模型,比如这个4000亿参数级别的模型,会在2026年第一季度这个时间线推出。
So I think this kind of balance is set to change this year in terms of
所以我认为这种平衡今年将会发生变化,就
what people are using the Chinese versus U.S. open models for,
人们使用中国与美国开源模型的用途而言,
which I'm personally going to be very excited to watch.
这一点我个人会非常期待看到。
First of all, huge props for being able to name so many of these.
首先,能说出这么多模型的名字,真是了不起。
huge props 地道口语
极大的赞赏、致敬
口语中表示对某人的高度认可,props 即 respect。
Did you actually name Llama?
Llama 真的是你命名的吗?
No. I feel like RIP.
不是。我感觉像是安息吧。
This was not on purpose.
这不是故意的。
on purpose 常用搭配
故意地
表示有意为之,常用于否认或强调意图。
RIP Llama.
安息吧,Llama。
All right. Can you mention what are some interesting models that stand out
好的。你能说说有哪些有趣的模型比较突出吗?
so your mission quen 3s is obviously a standout
所以你们的 Mission Quen 3s 显然很突出,
so I would say the years almost book ended by both DeepSeek version 3 and R1
所以我会说这一年几乎被 DeepSeek 的 V3 和 R1 首尾呼应,
and then on the other hand in December DeepSeek version 3.2
然后另一方面,在十二月 DeepSeek 3.2 版本
because what I like about those is they always have an interesting architecture tweak that others don't have
因为我喜欢它们的地方在于,它们总是有一些别人没有的有趣架构调整,
but otherwise if you want to go with um you know like the familiar but really good performance
但除此之外,如果你想用那种你熟悉的、但性能非常好的模型,
quen 3 and like um Nathan said also GPT OSS and I think GPT OSS
Quen 3,还有像 Nathan 说的 GPT OSS,而且我觉得 GPT OSS
what's interesting about it is kind of like the first public or like open weight model that was really trained with tool use in mind
它有趣的地方在于,它算是第一个真正以工具使用为核心训练的开源权重模型,
which I do think is kind of a bit of a paradigm shift where the ecosystem
我确实认为这算是一种范式转变,整个生态系统
paradigm shift 常用搭配
范式转变
指根本性的思维或方法转变,多用于科技、学术语境。
was not quite ready for it, so with tool use, I mean,
还没完全准备好,所以说到工具使用,我的意思是,
that the LLM is able to do a web search, to call a Python interpreter, and I do
大语言模型能够进行网络搜索、调用Python解释器,而我确实
think this, it's a standout, because I think it's a huge unlock,
认为这一点很突出,因为我觉得这是一个巨大的突破,
because one of the most common complaints about LLMs are, for example, hallucinations, right?
因为关于大语言模型最常见的抱怨之一就是,比如说,幻觉,对吧?
And so in my opinion, one of the best ways to solve hallucinations is to not try to always remember information or make things up.
所以在我看来,解决幻觉最好的方法之一就是不要总是试图记住信息或者编造内容。
For math, why not use a calculator app or Python?
对于数学,为什么不用计算器应用或者Python呢?
If I asked the LLM who won the soccer World Cup in 1998, instead of just trying to memorize, it could go do a search.
如果我问大语言模型1998年谁赢得了足球世界杯,与其只是试图记住,它可以去搜索一下。
I think mostly it's usually still Google search.
我觉得大多数情况下通常还是谷歌搜索。
So GPT, GPT OSS, they would do a tool call to Google, maybe find the FIFA website, find, okay, it was France.
所以GPT,GPT OSS,它们会调用工具去谷歌,也许找到国际足联网站,找到,好的,是法国。
do a tool call 常用搭配
调用工具(让模型去使用外部工具)
谈论AI模型调用外部工具或API时使用,技术讨论中常见。
It would get you that information reliably instead of just trying to memorize it.
它会可靠地给你那个信息,而不是只是试图记住它。
instead of just trying to memorize it 句型
而不是只是试图记住它
instead of just [doing something]
用于对比两种做法,强调应选择更可靠的方式而非死记硬背。
So I think it's a huge unlock, which I think right now is not fully utilized yet by the open source, open-weight ecosystem.
所以我觉得这是一个巨大的突破,而我认为目前开源、开放权重生态系统还没有充分利用它。
a huge unlock 常用搭配
一个巨大的突破/解锁
口语中形容某事物带来重大突破或释放巨大潜力。
A lot of people don't use tool call modes because I think it's first, it's a trust thing.
很多人不使用工具调用模式,因为我觉得首先这是一个信任问题。
it's a trust thing 地道口语
这是个信任问题
口语中解释某事的根本原因是信任,而非技术或其他因素。
You don't want to run this on your computer
你不想在你的电脑上运行这个
where it has access to tools could wipe your hard drive or whatever
因为它能访问工具,可能会擦除你的硬盘之类的
so you want to maybe containerize that
所以你可能想把它容器化
but I do think, you know, that that is like a really important step for the upcoming years to have this ability
但我确实认为,你知道,那真的是未来几年拥有这种能力的重要一步
so uh a few quick things first of all
所以呃,首先几件快事
first of all 地道口语
首先
口语中用于引出第一点,常用于列举或开启话题。
thank you for defining what you mean by tool
谢谢你定义了你所说的工具是什么意思
use I think that's a great thing to do in general for the concepts we're talking about
使用,我认为总的来说,对于我们正在谈论的概念,这是一件很棒的事情
even things as sort of well established as moes
即使是像 moes 这样已经相当成熟的东西
uh you have to say that means mixture of experts and you kind of have to build up an intuition for people,
呃,你必须说那意味着专家混合,而且你有点必须为人们建立一种直觉,
build up an intuition 常用搭配
建立直觉
用于描述通过积累逐渐形成对某事物的直觉理解。
what that means, how it's actually utilized, what are the different flavors.
那意味着什么,它实际上是如何被利用的,有哪些不同的风味。
So what does it mean that there's just such explosion of open models?
那么,开放模型如此爆炸式增长意味着什么?
What's your intuition? If you're releasing an open model, you want people to use it as the first and foremost thing.
你的直觉是什么?如果你发布一个开放模型,你希望人们首先使用它。
first and foremost 常用搭配
首要的,最重要的
强调某事物是最重要或最优先的,常用于正式或半正式场合。
And then after that comes things like transparency and trust.
然后在那之后才是透明度和信任之类的事情。
I think when you look at China, the biggest reason is that they want people around the world to use these models.
我认为当你看看中国,最大的原因是他们希望世界各地的人们使用这些模型。
And I think a lot of people will not, if you look outside of the U.S.,
而且我认为,如果你看看美国以外的地方,很多人不会,
a lot of people will not pay for software,
很多人不会为软件付费,
but they might have computing resources where you can put a model on it and run it.
但他们可能有计算资源,可以把模型放上去运行。
I think there can also be data that you don't want to send to the cloud.
我认为也可能有些数据你不想发送到云端。
So the number one thing is getting people to use models,
所以第一件事就是让人们使用模型,
use AI or use your AI that might not be able to do it without having access to the model.
使用AI,或者使用你的AI,而如果没有访问模型的权限,可能就无法做到。
I guess we should state explicitly.
我想我们应该明确说明。
So we've been talking about these Chinese models and open weight models.
所以我们一直在讨论这些中国模型和开放权重模型。
Oftentimes the way they're run is locally.
通常它们的运行方式是本地运行。
So it's not like you're sending your data to China or to whoever developed, to Silicon Valley, whoever developed the model.
所以并不是说你在把数据发送到中国,或者发送给开发者,发送到硅谷,发送给开发这个模型的人。
A lot of American startups make money by hosting these models from China and selling them, selling tokens.
很多美国初创公司通过托管这些来自中国的模型并出售它们、出售token来赚钱。
It's called like selling tokens, which means somebody will call the model to do some piece of work.
这叫做出售token,意思是有人会调用模型来完成某项工作。
I think the other reason is for U.S. companies, opening AI is so GPU deprived.
我认为另一个原因是对美国公司来说,开放AI非常缺乏GPU。
They're at the limits of the GPUs.
他们已经达到了GPU的极限。
Whenever they make a release, they're always talking about, like, our GPUs are hurting.
每当他们发布新版本时,他们总是说,我们的GPU很吃紧。
And I think there's, like, in one of these, like, GPT OSS release sessions,
而且我觉得,就像,在这些GPT OSS发布活动中,
Sam Altman said, like, oh, we're releasing this because we can use your GPUs.
Sam Altman说,就像,哦,我们发布这个是因为我们可以用你们的GPU。
We don't have to use our GPUs.
我们不必用我们自己的GPU。
And OpenAI can still get distribution out of this, which is another very real thing.
而OpenAI仍然可以从中获得分发,这是另一个非常真实的事情。
It doesn't cost them, though, anything.
不过,这并不花费他们任何东西。
And for the user, I think also, I mean, there are users who just use the model locally
而对于用户来说,我觉得也是,我的意思是,有些用户只是在本地使用模型,
how they would use GPD.
就像他们会用GPD那样。
but also for companies, I think it's a huge unlock to have these models
但对企业来说,我认为拥有这些模型是一个巨大的解锁,
because you can customize them, you can train them, you can add post-training, add more data, specialize them into, let's say, law, medical models, whatever you have.
因为你可以定制它们,你可以训练它们,你可以添加后训练,添加更多数据,将它们专门化为,比如说,法律、医疗模型,无论你有什么。
And the appeal, you mentioned LAMA,
而吸引力,你提到了LAMA,
the appeal of the open-weight models from China is that the open-weight models are also, the licenses are even friendlier.
来自中国的开放权重模型的吸引力在于,这些开放权重模型也是,许可证甚至更友好。
I think they are just unrestricted open-source licenses where if we use something like LAMA or JAMA, there are some strings attached.
我认为它们只是无限制的开源许可证,而如果我们使用像LAMA或JAMA这样的东西,就会有一些附加条件。
strings attached 常用搭配
附加条件
常用于描述协议、许可或交易中隐藏的限制或条件,口语和书面均可。
I think it's like upper limit in terms of how many users you have
我觉得这就像是你拥有多少用户的上限
and then if you exceed, I don't know, so
然后如果你超过了,我不知道,所以
and so many million users you have to report your finance situation to
而且你有那么多百万用户,你就必须向……报告你的财务状况
let's say matter or something like that
比如说Matter之类的
and I think well it is a free model
然后我觉得,嗯,它确实是个免费模型
but there are strings attached
但是是有附加条件的
there are strings attached 常用搭配
有附加条件
用于说明某事并非完全无条件,存在隐含的限制或要求。
and people do like things where strings are not attached
而人们确实喜欢没有附加条件的东西
so I think that's also one of the reasons besides performance
所以我觉得这也是除了性能之外的其中一个原因
why the open weight models from China are so popular
为什么来自中国的开放权重模型这么受欢迎
because you you can just use them there's no catch in that sense.
因为你可以直接使用它们,从这个意义上说没有任何陷阱。
there's no catch 地道口语
没有任何陷阱/隐藏条件
口语中表示某事没有隐藏的代价或问题,可以放心接受。
The ecosystem has gotten better on that front,
在这方面生态系统已经变得更好了,
but mostly downstream of these new providers providing such open licenses.
但主要是这些提供此类开放许可证的新供应商的下游。
That was funny when you pulled up Perplexity and said Kimi K2 Thinking hosted in the US,
当你打开Perplexity并说Kimi K2 Thinking托管在美国时,那很有趣,
which is just like an exact, I've never seen this, but it's an exact example of what we're talking about where people are sensitive to this.
这就像是一个确切的,我从未见过这个,但它是我们正在谈论的人们对此敏感的一个确切例子。
Like Kimi K2 Thinking and Kimi K2 is a model that is very popular.
就像Kimi K2 Thinking和Kimi K2是一个非常受欢迎的模型。
People say that has very good creative writing and also in doing some software things.
人们说它在创意写作以及做一些软件方面非常出色。
It's just these little quirks that people pick up on with different models that they like.
这只是人们在不同模型上发现的、他们喜欢的一些小怪癖。
pick up on 常用搭配
注意到,察觉到
指注意到别人可能忽略的细节或特点,常用于口语。
What are some interesting ideas that some of these models have explored that you can speak to, like that particularly interesting to you?
这些模型探索过哪些有趣的思路,是你可以谈谈的,比如你觉得特别有意思的?
Maybe you can go chronologically.
也许你可以按时间顺序来讲。
I mean, there was, of course, DeepSeq R1 that came out in January if we just focus on 2025.
我是说,当然有DeepSeq R1,如果只聚焦2025年的话,它是一月发布的。
However, this was based on DeepSeq version 3, which came out the year before in December 2024.
然而,这是基于DeepSeq版本3,它在前一年2024年12月发布。
There are multiple things on the architecture side.
在架构方面有多个东西。
What is fascinating is you can still, I mean, that's what I do in my from scratch coding projects.
有趣的是,你仍然可以,我是说,这就是我在从零开始编码项目里做的事。
from scratch 常用搭配
从零开始
指从头开始做某事,不借助已有的基础或材料。
you can still start with GPT-2 and you can add things to that model to make it into this other model.
你仍然可以从GPT-2开始,然后你可以给那个模型添加东西,把它变成另一个模型。
So it's all still kind of like the same lineage, the same, it is a very close relationship between those.
所以它仍然有点像同一个谱系,同样的,它们之间关系非常紧密。
But top of my head, DeepSeq, what was unique there is the mixture of experts.
但凭记忆,DeepSeq,独特之处在于专家混合。
top of my head 地道口语
凭记忆,临时想到
口语中表示没有准备、仅凭记忆说出,常用于回答问题时。
I mean, they were not inventing mixture of experts.
我是说,他们并没有发明专家混合。
We can maybe talk a bit more what mixture of experts means.
我们也许可以多谈谈专家混合是什么意思。
But just to list these things first, before we dive into detail, mixture of experts,
但先列出这些东西,在我们深入细节之前,专家混合,
But then they also had multi-head latent attention, which is a tweak to the attention mechanism, where this was, I would say, 2025, the main distinguishing factor between these open weight models.
但然后他们还有多头潜在注意力,这是对注意力机制的一个调整,我会说,这是2025年这些开放权重模型之间的主要区别因素。
Different tweaks to make inference or KV cache size.
不同的调整,用来改变推理或KV缓存大小。
We can also define KV cache in a few moments,
我们稍后也可以定义一下KV缓存,
but to kind of make it more economical, to have long contacts, to shrink the KV cache size.
但为了让它更经济,能处理长上下文,来缩小KV缓存大小。
So what are tweaks that we can do?
那么我们可以做哪些调整呢?
And most of them focused on the attention mechanism.
其中大多数都集中在注意力机制上。
focused on 常用搭配
集中在……上
用于说明注意力、讨论或工作的重点所在。
There is multi-hat latent attention in DeepSeq.
DeepSeq中有多头潜在注意力。
There is group query attention, which is still very popular.
还有分组查询注意力,它仍然非常流行。
It's not invented by any of those models.
它不是由那些模型中的任何一个发明的。
It goes back a few years, but that would be the other option.
它可以追溯到几年前,但那会是另一个选项。
goes back a few years 常用搭配
可以追溯到几年前
说明某事物的起源或历史并不新,常用于澄清发明时间。
Sliding window attention, I think, almost reuses it, if I remember correctly.
滑动窗口注意力,我想,如果我没记错的话,几乎复用了它。
if I remember correctly 地道口语
如果我没记错的话
口语中用于对自己所说内容表示不确定,留有余地。
So there are these different tweaks that make the models different.
所以有这些不同的调整,让模型变得不同。
Otherwise, I put them all together in an article once where I just compared them.
否则,我曾经把它们都放在一篇文章里,我只是比较了它们。
put them all together 常用搭配
把它们都放在一起
用于描述将多个事物汇总或集中到一处。
are very surprisingly similar it's just different numbers in terms of how many repetitions of the transformer block you have in the center
非常惊人地相似,只是在中心有多少次transformer块重复方面数字不同
in terms of 常用搭配
在……方面
用于引出讨论的具体方面或角度。
and like just little little knobs that people tune but what what's so nice about it is it's it works no matter what you can tweak things you can
还有人们调整的那些小旋钮,但它好就好在,无论你怎么调整,它都能工作
what's so nice about it is 句型
它好就好在……
what's so nice about [something] is [clause]
用于强调某事物最令人满意的优点。
move the normalization layers around you get some performance gains and i almost always very good in ablation studies
移动归一化层的位置,你会获得一些性能提升,而且我几乎总是在消融研究中表现很好
move the normalization layers around 常用搭配
调整归一化层的位置
描述对模型结构做位置上的调整。
showing what actually what it does to the model if you move something around ablation studies doesn't
展示如果你移动某些东西,实际上对模型有什么影响,消融研究并不
make it better or worse,
让它变得更好或更糟,
but there are so many
但有很多
let's say ways you can implement a transformer and make it still work.
比如说,你可以用很多方式实现一个Transformer并让它仍然有效。
let's say 地道口语
比如说
口语中用于举例或提出假设情况。
Big ideas that are still prevalent is mixture of experts,
仍然流行的大思想是专家混合,
multi-latent attention, sliding window attention, group query attention,
多潜在注意力、滑动窗口注意力、分组查询注意力,
and then at the end of the year we saw a focus on making the attention mechanism scale linearly with inference token prediction.
然后年底我们看到一个重点,就是让注意力机制随推理token预测线性扩展。
at the end of the year 常用搭配
在年底
用于指一年接近结束的时间段。
So there were quen3next for example, which added a gated deltanet.
所以有比如quen3next,它加入了一个门控deltanet。
It's kind of like inspired by state-based models,
它有点像受基于状态的模型启发,
kind of like 地道口语
有点像
口语中用于模糊地描述相似性,语气较随意。
where you have a fixed state that you keep updating,
你有一个固定状态并不断更新,
but it makes essentially this attention cheaper, or it replaces attention with a cheaper operation.
但它本质上让这种注意力更便宜,或者用更便宜的操作替代注意力。
And it maybe is it useful to step back and talk about transform architecture in general?
也许退一步谈谈Transformer架构总体上有用吗?
step back and talk about 常用搭配
退一步谈谈
用于建议从更宏观或更基础的角度讨论问题。
Yeah so maybe we should start with the GPT-2 architecture, the transformer that was derived from the Attention Is All You Need paper.
是的,所以也许我们应该从GPT-2架构开始,这个Transformer源自《Attention Is All You Need》论文。
start with 常用搭配
从……开始
用于建议以某事物作为讨论或过程的起点。
So the Attention Is All You Need paper had a transformer architecture that had two parts, an encoder and a decoder,
所以《Attention Is All You Need》论文有一个Transformer架构,它有两部分:编码器和解码器,
and GPT went just focusing in on the decoder part.
而GPT只专注于解码器部分。
focusing in on 常用搭配
专注于
用于说明把注意力集中到某个具体部分。
It is essentially still a neural network and it has this attention mechanism inside,
它本质上仍然是一个神经网络,并且内部有这个注意力机制,
and you predict one token at a time, you pass it through an
你一次预测一个token,你把它通过一个
embedding layer. There's the transformer block.
嵌入层。然后是 Transformer 模块。
The transformer block has attention modules and a fully connected layer, and there are some normalization layers in between.
Transformer 模块包含注意力模块和一个全连接层,中间还有一些归一化层。
But it's essentially neural network layers with this attention mechanism.
但它本质上就是带有这种注意力机制的神经网络层。
So coming from GPT-2, when we move on to GPT-OSS,
所以从 GPT-2 到 GPT-OSS,
move on to 常用搭配
转向;接着讨论
用于从当前话题过渡到下一个话题。
there is, for example, the mixture of experts layer.
比如说,有混合专家层。
It's not invented by GPT-OSS. It's a few years old.
它不是 GPT-OSS 发明的。它已经有好几年历史了。
But it is essentially a tweak to make the model larger without consuming more compute in each forward pass.
但它本质上是一种调整,让模型变得更大,而不会在每次前向传播中消耗更多算力。
So there is this fully connected layer.
所以有一个这样的全连接层。
And if listeners are familiar with multilayer perceptrons, you can think of a mini multilayer perceptron, a fully connected neural network layer inside the transformer.
如果听众熟悉多层感知机,你可以把它想象成 Transformer 内部的一个小型多层感知机,也就是一个全连接神经网络层。
think of 常用搭配
把……想象成
用于帮助理解抽象概念,给出类比。
And it's very expensive because it's fully connected.
而且它非常昂贵,因为它是全连接的。
If you have a thousand inputs, a thousand outputs, it's like one million connections.
如果你有一千个输入、一千个输出,那就相当于一百万个连接。
And it's a very expensive part in this transformer.
它是这个 Transformer 中非常昂贵的一部分。
And the idea is to kind of expand that into multiple feedforward networks.
而思路就是把它扩展成多个前馈网络。
So instead of having one, let's say you have 256, but it would make it way more expensive
所以与其只有一个,假设你有 256 个,但那会让它昂贵得多,
instead of 常用搭配
而不是
用于对比两种选择,说明用前者替代后者。
because now you have 256, but you don't use all of them at the same time.
因为现在你有 256 个,但你不会同时使用全部。
at the same time 常用搭配
同时
用于说明两件事同时发生或同时成立。
So you now have a router that says, okay, based on this input token, it would be useful to use this fully connected network.
所以你现在有一个路由器,它会说,好的,基于这个输入词元,使用这个全连接网络会很有用。
And in that context, it's called an expert.
而在那种语境下,它被称为专家。
So a mixture of experts means you have multiple experts.
所以混合专家意味着你有多个专家。
And depending on what your input is, let's say it's more math heavy, it would use different experts compared to, let's say, translating input text from English to Spanish.
而根据你的输入是什么,比如说它更偏重数学,它会使用不同的专家,相比于,比如说,将输入文本从英语翻译成西班牙语。
depending on 常用搭配
取决于
用于说明结果会随某条件而变化。
It would maybe consult different experts.
它可能会咨询不同的专家。
It's not quite clear. I mean, it's clear cut to say, OK, this is only an expert for math and for Spanish is a bit more fuzzy.
这不太清楚。我的意思是,明确地说,好吧,这只是一个数学专家,而对于西班牙语则有点模糊。
I mean 地道口语
我的意思是
口语中用于补充、澄清或修正前面的话。
But the idea is essentially that you pack more knowledge into the network, but not all the knowledge is used all the time.
但本质上,这个想法是你将更多知识打包进网络,但并非所有知识都会一直被使用。
That would be very wasteful.
那会非常浪费。
So you're kind of like during the token generation, you're more selective.
所以你在词元生成过程中,有点像是更有选择性地使用。
kind of like 地道口语
有点像
口语中用于模糊地描述相似性,语气较随意。
There's a router that selects which tokens should go to which expert.
有一个路由器会选择哪些词元应该去哪个专家。
It's more complexity. It's harder to train.
这更复杂。也更难训练。
There's a lot of, you know, that can go wrong, like collapse and everything.
有很多,你知道,可能出错的地方,比如崩溃之类的。
go wrong 常用搭配
出问题;出错
用于说明事情可能失败或出现意外情况。
So I think that's why Olmo 3 still uses dense.
所以我认为这就是为什么Olmo 3仍然使用稠密模型。
I mean, you have, I think, Olmo models with mixture of experts, but dense models where dense means also it's jargon.
我的意思是,你有,我认为,Olmo模型有混合专家,但也有稠密模型,这里的稠密也是术语。
There's a distinction between dense and sparse.
稠密和稀疏之间有一个区别。
So a mixture of experts is considered sparse because we have a lot of experts, but only few of them are active.
所以专家混合被认为是稀疏的,因为我们有很多专家,但只有少数是活跃的。
So that's called sparse.
所以这被称为稀疏。
And then dense would be the opposite where you only have like one fully connected module and it's always, you know, utilized.
然后密集则是相反的,你只有一个全连接模块,而且它总是被使用。
So maybe this is a good place to also talk about KVCache, but actually before that, even zooming out, like fundamentally, how many new ideas have been implemented from GPT-2 to today?
所以也许这是一个好地方来谈谈KVCache,但实际上在那之前,甚至从更宏观的角度来看,从根本上说,从GPT-2到今天实施了多少新想法?
Like how different really are these architectures?
比如这些架构到底有多大不同?
picture like the mixture of experts um the attention mechanism in gpt oss that would be the group query attention mechanism so it's a slight tweak for multi-head attention to group query attention so that we have two
比如专家混合,嗯,GPT OSS中的注意力机制,那将是分组查询注意力机制,所以这是对多头注意力的一点小调整,变成分组查询注意力,这样我们有两个
i think they replaced a layer norm by rms norm but it's just like a different normalization layer and not a big change it's just like a tweak um the non-linear activation function people familiar in with deep new networks i mean it's the same as changing sigmoid with relu it's it's not changing the network fundamentally it's just like a tweak like a little little tweak um and that's about it i would say it's not really fundamentally that different it's still the same same architecture so you can convert one from one uh you can
我认为他们用RMS norm替换了层归一化,但这只是一个不同的归一化层,不是什么大变化,只是一个小调整。嗯,非线性激活函数,熟悉深度神经网络的人,我的意思是,这就像把sigmoid换成relu一样,它并没有从根本上改变网络,只是一个小调整,一点点小调整。嗯,差不多就是这样。我会说它并没有根本上的不同,仍然是相同的架构,所以你可以从一个转换到另一个,嗯,你可以
that's about it 地道口语
差不多就是这样
口语中用于表示列举或说明基本结束。
go from one into the other by just adding these, these changes, basically.
基本上,只要加上这些改动,就能从一个变成另一个。
It's fundamentally is still the same architecture.
它从根本上说还是同样的架构。
Yep. So for example, you mentioned my book earlier.
对。比如说,你之前提到过我的书。
That's a GPD2 model in the book because it's simple and it's very small, um, so 124, 120 million parameters approximately.
书里那个是GPD2模型,因为它简单而且非常小,嗯,大约1.24亿、1.2亿个参数。
But in the bonus materials, I do have almost three from-scratch Gemma 3, from-scratch, and other types of from-scratch models.
但在附加材料里,我确实有差不多三个从零开始训练的Gemma 3,从零开始,还有其他类型的从零开始训练的模型。
And I always started with my GPD2 model and just, you know, tweaked a few added different components.
而我一直都是从我的GPD2模型开始,然后就是,你知道,调整一下,加几个不同的组件。
And you get from one to the other, it's like, it's kind of like a lineage in a sense.
然后你就从一个得到了另一个,这有点像,某种意义上像是一条谱系。
in a sense 常用搭配
在某种意义上
用于表示从某个角度看某事成立,但并非完全字面。
Yeah, can you build up an intuition for people?
对,你能帮大家建立一种直觉吗?
build up an intuition 常用搭配
建立直觉
用于帮助他人通过解释逐步形成对某事物的直观理解。
Because sort of when you zoom out, you look at it, there's so much rapid advancement in the AI world.
因为当你拉远来看,你会发现,AI领域的进步是如此迅速。
zoom out 常用搭配
拉远来看;从宏观角度
用于建议从更宏观的视角审视问题。
At the same time, fundamentally, the architectures have not changed.
但与此同时,从根本上说,架构并没有改变。
So where is all the turbulence, the turmoil of the advancement happening?
那么,进步带来的所有这些动荡、混乱都发生在哪里呢?
Where, where's the gains to be had?
哪里,哪里能获得收益呢?
So there are the different stages where you develop the network or train the network.
所以有不同的阶段,你在这些阶段开发网络或训练网络。
You have the pre-training now, um, back then they, it was just pre-training with GPT2.
现在你有预训练,嗯,那时候他们,就只是用GPT2做预训练。
Now you have pre-training, mid-training, and post-training. Um, so I think right...
现在你有预训练、中期训练和后训练。嗯,所以我觉得,对……
Now we are in the post-training focus stage.
现在我们处于后训练重点阶段。
I mean, pre-training still gives you advantages if you scale it up to better, higher quality data,
我的意思是,预训练仍然会给你带来优势,如果你把它扩展到更好、更高质量的数据,
scale it up 常用搭配
扩大规模
用于描述把系统、项目或数据等做大、扩展。
but then we have capability unlocks that were not there with GPT-2.
但随后我们有了GPT-2所没有的能力解锁。
For example, ChatGPT, it is basically a GPT-3 model,
例如,ChatGPT,它基本上是一个GPT-3模型,
and GPT-3 is the same as GPT-2 in terms of architecture.
而GPT-3在架构上与GPT-2相同。
in terms of 常用搭配
在……方面
用于引出讨论的某个具体方面或角度。
What was new was adding the supervised fine-tuning and the reinforcement learning with human feedback.
新的是加入了监督微调和基于人类反馈的强化学习。
So it's more on the algorithmic side rather than the architecture
所以这更多是在算法方面,而不是架构方面
I would say that the systems also change a lot
我想说系统也变化很大
I think if you listen to Nvidia's announcements
我认为如果你听英伟达的公告
they talk about these things like you now do FP8, you can now do FP4
他们谈论这些事情,比如你现在可以做FP8,你现在可以做FP4
and what is happening is these labs are figuring out how to utilize more compute to put it into one model
正在发生的是,这些实验室正在弄清楚如何利用更多计算资源将其投入一个模型
figuring out 常用搭配
弄清楚、搞明白
表示通过思考或尝试找到解决办法,常用于口语。
which lets them train faster and that lets them put more data in and then you can find better configurations faster by doing this
这让他们训练得更快,也让他们能放入更多数据,然后通过这样做你可以更快地找到更好的配置
so you can look at like the essentially the tokens per second per GPU is a metric that you look at when you're doing large-scale training
所以你可以看,基本上每GPU每秒的token数是你进行大规模训练时关注的一个指标
and you could get you can go from like 10k to 13k by turning on FP8 training which
你可以从大约1万提升到1.3万,通过开启FP8训练,这
means they're using less memory per parameter in the model
意味着它们在模型中每个参数使用的内存更少
and by saving less information you do less communication you can train faster
而且通过保存更少的信息,你进行的通信更少,你可以训练得更快
so all these like system things underpin way faster experimentation on data and algorithms
所以所有这些系统层面的东西支撑着在数据和算法上快得多的实验
that is kind of like it's this it's this kind of loop that keeps going
这有点像,就是这种不断循环的过程
where it's kind of hard to describe when you look at the architecture
当你看着架构的时候,这有点难以描述
and they're exactly the same but the code base used to train these models is going to be vastly different
它们完全一样,但用来训练这些模型的代码库会大不相同
and you could probably like I don't the GPUs are different
而且你可能,我不知道,GPU 是不同的
but you probably train GPT-OSS 20B way faster in wall clock time than GPT-2 was trained at the time
但你可能训练 GPT-OSS 20B 的实际时间比当时训练 GPT-2 要快得多
yeah like you said they had for example in the mixture of experts
是的,就像你说的,他们在混合专家模型中有比如
this NVFP4 optimization for example where you get more throughput
比如这个 NVFP4 优化,你可以获得更高的吞吐量
but I do think this is for the speed this is true
但我确实认为这是为了速度,这是真的
but uh it doesn't give the model new capabilities in a sense
但呃,从某种意义上说,它并没有给模型带来新的能力
in a sense 常用搭配
在某种意义上
用于表示从某个角度看是这样,常带有保留或限定语气。
it's just how much can we make the computation coarser without suffering in terms of model performance degradation.
它只是我们能在不让模型性能下降的情况下把计算做得多粗糙。
But I do think, I mean, there are alternatives popping up to the transformer.
但我确实认为,我的意思是,正在出现一些替代 Transformer 的方案。
popping up 常用搭配
不断出现、冒出来
用于描述新事物或新方案开始出现,口语常用。
There's text diffusion models, completely different paradigm.
有文本扩散模型,完全不同的范式。
And there's also, I mean, text diffusion models might use transformer architectures,
而且还有,我是说,文本扩散模型可能会使用Transformer架构,
but it's not an auto-regressive transformer.
但它不是自回归Transformer。
And also Mamba models, it's a state-space model, but they do have trade-offs.
还有Mamba模型,它是一种状态空间模型,但它们确实有取舍。
trade-offs 常用搭配
权衡、取舍
指为了得到某方面的好处而必须接受另一方面的损失。
And what's right is there's nothing that has replaced the autoregressive transformer as state-of-the-art model.
而正确的是,没有什么能取代自回归Transformer作为最先进的模型。
So like for state-of-the-art, you would still do that, go with that thing.
所以对于最先进的,你仍然会那样做,选择那个东西。
But there are no alternatives for the cheaper and like alternatives that are kind of making compromises.
但对于更便宜的替代方案,以及那些做出妥协的替代方案,并没有其他选择。
But it's not just one architecture anymore.
但现在已经不只有一种架构了。
There are little ones coming up.
有一些小的正在出现。
But if we talk about the state-of-the-art, it's pretty much still the transformer architecture autoregressive derived from GPT-2 essentially.
但如果我们谈论最先进的,基本上仍然是从GPT-2衍生出的自回归Transformer架构。
I guess the big question here is we talked quite a bit here on the architecture behind the pre-training.
我想这里的大问题是,我们在这里已经谈了很多关于预训练背后的架构。
Are the scaling laws holding strong across pre-training, post-training, inference, context size, data, synthetic data?
缩放定律在预训练、后训练、推理、上下文大小、数据、合成数据方面是否依然成立?
I like to start with the technical definition of scaling law, which kind of informs all of this.
我喜欢从缩放定律的技术定义开始,这在一定程度上为所有这些提供了信息。
The scaling law is a power law relationship between you can think of the x-axis.
缩放定律是一种幂律关系,你可以把它看作x轴。
So kind of what you are scaling is a combination of compute and data, which are kind of similar.
所以你所缩放的是计算和数据的组合,它们有点相似。
And then the y-axis is like the held out prediction accuracy over our next tokens.
然后y轴就像是对我们下一个token的留出预测准确率。
We talk about models being auto-regressive.
我们谈论模型是自回归的。
It's like if you keep a set of text that the model has not seen, how accurate will it get when you will train?
就像如果你保留一组模型没见过的文本,当你训练时它会有多准确?
And the idea of scaling laws came when people figured out that that was a very predictable relationship.
而缩放定律的概念出现于人们发现那是一种非常可预测的关系时。
And I think that that technical term is continuing.
我认为那个技术术语是持续性的。
And then the question is like, what do users get out of it?
然后问题就像是,用户能从中得到什么?
And then there are more types of scaling where OpenAI's 01 was famous for introducing inference time scaling.
然后还有更多类型的缩放,其中 OpenAI 的 01 因引入推理时间缩放而闻名。
And I think less famously for also showing that you can scale reinforcement learning training and get kind of this log X axis and then a linear increase in performance on Y axis.
我认为不太出名的是,它还展示了你可以扩展强化学习训练,并得到这种对数 X 轴,然后在 Y 轴上性能线性增长。
So there's kind of these three axes now where the traditional scaling laws are talked about for pre-training,
所以现在有这三种轴,其中传统的缩放定律是针对预训练的,
which is how big your model is and how big your data set is.
即你的模型有多大以及你的数据集有多大。
And then scaling reinforcement learning, which is like, how long can you do this trial and error learning that we will talk about?
然后是扩展强化学习,这就像,你能进行这种我们将要讨论的试错学习多久?
We'll define more of this.
我们会进一步定义这一点。
And then this inference time compute, which is just letting the model generate more tokens on a specific problem.
然后是推理时间计算,这只是让模型在特定问题上生成更多 token。
So I'm kind of bullish where they're all really still working.
所以我有点看涨,它们都真的还在起作用。
But the low hanging fruit has mostly been taken, especially in the last year
但低垂的果实大多已经被摘走了,尤其是在过去一年里
low hanging fruit 地道口语
容易摘取的成果、容易实现的目标
比喻最容易获得或最容易完成的部分,常用于讨论机会或进展。
on reinforcement learning with verifiable rewards, which is this RLVR.
在带可验证奖励的强化学习上,也就是这个 RLVR。
And then inference time scaling, which is just why these models feel so different to use
然后是推理时间扩展,这正是为什么这些模型用起来感觉如此不同
where previously you would get that first token immediately.
以前你会立刻得到第一个词元。
And now they'll go off for seconds, minutes, or even hours generating these hidden thoughts before giving you the first word of your answer.
而现在它们会花上几秒、几分钟甚至几小时来生成这些隐藏的思考,然后才给出你答案的第一个词。
And that's all about this inference time scaling, which is such a wonderful kind of step function in terms of how the models change abilities.
而这一切都关乎这个推理时间扩展,它在模型能力变化方面就像一种奇妙的阶跃函数。
They kind of enabled this tool use stuff and enabled this much better software engineering that we were talking about.
它们某种程度上实现了这种工具使用,也实现了我们之前谈到的这种好得多的软件工程。
And this is when we say enabled almost entirely downstream of the fact that this reinforcement learning with verifiable rewards training just kind of let the models pick up these skills very easily.
这就是为什么我们说,几乎完全是因为这种带可验证奖励的强化学习训练,让模型非常容易地掌握了这些技能。
So let the models learn.
所以让模型去学习吧。
So if you look at the reasoning process when the models are generating a lot of tokens,
所以如果你观察模型生成大量词元时的推理过程,
what it will be often doing is it tries a tool.
它经常会做的是尝试一个工具。
It looks at what it gets back.
它看看返回了什么。
It tries another API.
它再试另一个 API。
It sees what it gets back and if it solves the problem.
它看看返回了什么,以及是否解决了问题。
So the models, when you're training them, very quickly learn to do this.
所以这些模型,在你训练它们的时候,很快就学会做这件事。
And then at the end of the day, that gives this kind of general foundation where the model can use CLI commands very nicely in your repo and handle Git for you and move things around and organize things or search to find more information.
然后最终,这提供了一种通用的基础,让模型可以在你的代码仓库中很好地使用命令行命令,替你处理 Git,移动东西、整理东西,或者搜索以找到更多信息。
at the end of the day 地道口语
归根结底、最终
用于总结最终结果或最重要的结论,口语常用。
Which if we're sitting in these chairs a year ago, it's something that we didn't really think of the models being doing.
如果一年前我们坐在这几把椅子上,这是我们当时真的没想到模型会做的事情。
So this is just kind of something that has happened this year and has totally transformed how we think of using AI, which I think is very magical.
所以这只是今年发生的一件事,它彻底改变了我们对使用 AI 的看法,我觉得这非常神奇。
such an interesting evolution and just so unlock so much value
如此有趣的演变,并且释放了如此多的价值
but it's like it's not clear what the next avenue will be in terms of unlocking stuff like this
但就像,不清楚下一个途径会是什么,在解锁这类东西方面
i think that there's there's we'll get to continual learning later but there's a lot of buzz around certain areas of ai but no one knows when the next step function will really come so
我认为,我们稍后会谈到持续学习,但 AI 的某些领域有很多热议,但没人知道下一个阶跃函数真正何时会到来,所以
you you've actually said quite a lot of things there and said profound things quickly it would be nice to unpack them a little bit
你实际上在那里说了很多,而且很快说出了深刻的东西,能稍微展开一下就好了
you said you're bullish basically on every version of scaling.
你说你基本上看好每一种扩展方式。
So we just even start at the beginning, pre-training,
所以我们甚至从一开始,预训练,
Are we kind of implying that the low-hanging fruit on pre-training scaling has been picked?
我们是不是在暗示预训练扩展的低垂果实已经被摘完了?
low-hanging fruit 常用搭配
最容易实现的目标或最容易取得的成果
用于描述某个领域中容易达成的部分,常与已被摘取/已用完搭配。
Has pre-training hit a plateau, or is even pre-training still you're bullish on?
预训练是已经进入平台期了,还是你依然看好预训练?
hit a plateau 常用搭配
达到平台期,停止增长
用于描述进展、增长或性能在一段时间后不再提升。
Pre-training has gotten extremely expensive.
预训练已经变得极其昂贵。
I think to scale up pre-training, it's also implying that you're going to serve a very large model to the users.
我认为要扩大预训练规模,这也意味着你要向用户提供一个非常大的模型。
So I think that it's been loosely established.
所以我认为这已经大致确立了。
the likes of GPT-4 and similar models were around this order of trillion parameters at the biggest size.
像GPT-4和类似的模型,在最大规模时大约是这个万亿参数的量级。
the likes of 常用搭配
像……这样的人或事物
用于举例,指与所提到的人或事物同类的对象。
There's a lot of rumors that they've actually gotten smaller as training has gotten more efficient.
有很多传言说,随着训练变得更高效,它们实际上变得更小了。
You want to make the model smaller because then your costs of serving go down proportionally.
你想让模型更小,因为这样你的服务成本就会成比例下降。
These models, the cost of training them is really low relative to the cost of serving them to hundreds of millions of users.
这些模型,训练它们的成本相对于向数亿用户提供服务的成本来说真的很低。
I think DeepSeek had this famous number of about $5 million for pre-training at cloud market rates,
我认为DeepSeek有一个著名的数字,按云市场价格计算,预训练大约花了500万美元,
I think Olmo 3, section 2.4 in the paper, we just detailed how long we had the GPU clusters sitting around for training,
我想Olmo 3,论文中的2.4节,我们详细说明了GPU集群闲置等待训练的时间有多长,
which includes engineering issues, multiple seeds, and it was like about $2 million to rent the cluster to deal with all the problems and headaches of training a model.
其中包括工程问题、多个随机种子,租用集群来处理训练模型的所有问题和麻烦大约花了200万美元。
So these models are pretty, like a lot of people could get $1 to $10 million to train a model,
所以这些模型相当……比如很多人可以花100万到1000万美元来训练一个模型,
but the recurring costs of serving millions of users is really billions of dollars of compute.
但服务数百万用户的经常性成本实际上是数十亿美元的计算资源。
I think that you can look at close, like a thousand GPU rental,
我觉得你可以看看,比如一千块GPU的租赁,
you can pay a hundred grand a day for, and these companies could have millions of GPUs.
你每天可能要付十万美元,而这些公司可能拥有数百万块GPU。
You can look at how much these things cost to sit around.
你可以看看这些东西闲置着要花多少钱。
So that's kind of a big thing.
所以这算是个大问题。
And then it's like, if scaling is actually giving you a better model, like, is it going to be financially worth it?
然后就像,如果扩大规模确实能给你一个更好的模型,那么,这在经济上值得吗?
And I think it'll kind of slowly, we'll push it out as AI solves more compelling tasks.
我觉得它会慢慢地,随着AI解决更多有吸引力的任务,我们会把它推出去。
So like the likes of Cloud Opus 4.5, making cloud code just work for things
所以像Cloud Opus 4.5这样的,让Cloud代码能直接用于各种事情
i think i launched this project called like the atom project
我想我启动了这个叫Atom项目的项目
which is like american truly open models in july
它就像是七月的美国真正开放模型
and that was like a true vibe coded website
那是一个真正凭感觉编码的网站
vibe coded 地道口语
凭感觉用AI辅助快速写代码做出来的
口语中形容不经过严格工程流程、靠直觉和AI工具快速搭建的项目。
and like i have a job um make plots and stuff
而且我有一份工作,嗯,做图表之类的
and then i came back to refresh it in the last few weeks
然后我在过去几周回来更新了它
and it's like cloud opus 4.5 versus whatever model at the time was like just crushed all the issues
结果就像Cloud Opus 4.5对比当时的其他模型,简直碾压了所有问题
that it had from building in june and
它在六月构建时遇到的那些问题,而且
July and like it might be a bigger model.
七月,而且它可能是一个更大的模型。
There's a lot of things that go into this, but that's like, there's still progress coming.
这里面涉及很多因素,但就像,仍然有进展在出现。
So what you're speaking to is the nuance of the y-axis of the scaling laws,
所以你说的是缩放定律纵轴上的细微差别,
that the way it's experienced versus on a benchmark, the actual intelligence might be different.
即它被体验的方式与在基准测试上的表现相比,实际智能可能会有所不同。
But still, your intuition about pre-training, if you scale the size of compute, will the models get better?
但仍然,你对预训练的直觉是,如果你扩大计算规模,模型会变得更好吗?
Not whether it's financially viable, but just from the law aspect of it, do you think the models will get smarter?
不是问它在财务上是否可行,而只是从定律方面来看,你认为模型会变得更聪明吗?
Yeah. And I think that there's and this sometimes comes off as like almost like disillusioned from people, leadership, AI companies saying this,
是的。而且我认为,有时候这听起来几乎像是来自人们、领导层、AI公司的一种幻灭感,他们这么说,
comes off as 常用搭配
给人某种印象,显得像是
用于描述某句话或行为给他人留下的印象,常接形容词或名词。
but they're like, it's held for 13 orders of magnitude of computers or something like why would it ever end?
但他们说,这已经保持了13个数量级的计算机规模之类的,那它为什么会结束呢?
So I think fundamentally it is pretty unlikely to stop.
所以我认为从根本上说,它不太可能停止。
It's just like eventually we're not even going to be able to test the bigger scales because of all the problems that come with more compute.
就像最终我们甚至无法测试更大的规模,因为更多计算带来的所有问题。
I think that there's a lot of talk on how 2026 is a year when very large Blackwell compute clusters, like gigawatt scale facilities, hyperscalers are coming online.
我认为有很多讨论关于2026年将是大型Blackwell计算集群,比如吉瓦级设施、超大规模企业上线的一年。
And these were all contracts for power and data centers that were signed and sought out in like 22 and 2023.
这些都是2022年和2023年左右签署和寻求的电力与数据中心合同。
So before or right after ChatGPT.
所以是在ChatGPT之前或之后不久。
So it took this two to three year lead time to build these bigger clusters to train the models.
所以建造这些更大的集群来训练模型需要两到三年的前置时间。
While there's obviously immense interest in building even more data centers than that.
虽然显然有极大的兴趣建造比那更多的数据中心。
So that is like kind of the crux that people are saying is like these new clusters are coming.
所以这就像是人们所说的关键点,这些新集群即将到来。
The labs are gonna have more compute for training.
实验室将拥有更多用于训练的算力。
They're going to utilize this, but it's not a given.
他们会利用这一点,但这并不是必然的。
it's not a given 常用搭配
这并非必然,不能想当然
用于强调某事没有保证,需要条件或努力才能实现。
And it's like, I've seen so much progress that I expect it.
而且就像,我看到了如此大的进步,所以我期待它。
And I expect a little bit bigger models.
我期待稍微大一点的模型。
And I expect, I would say it's more like we'll see a $2,000 subscription this year.
我预计,我会说更像是我们今年会看到2000美元的订阅。
We see $200 subscriptions. It's like that can 10x again.
我们看到200美元的订阅。就像那可以再增长10倍。
And these are the kinds of things that could come and they're all downstream of this like bit bigger model that offers just a little bit more cutting edge.
而这些是可能出现的事情,它们都是这个稍大一点的模型的下游产物,这个模型提供稍微更前沿一点的能力。
So, you know, it's reported that XAI is going to hit that one gigawatt scale early 26 and full two gigawatt by year end.
所以,你知道,据报道XAI将在26年初达到1吉瓦规模,并在年底达到完整的2吉瓦。
How do you think they'll utilize that in the context of scaling laws?
你认为他们会在缩放定律的背景下如何利用这一点?
Is a lot of that inference, is a lot of that training?
其中很多是推理,还是很多是训练?
It ends up being all of the above.
最终是以上所有。
all of the above 常用搭配
以上所有(选项或情况)
用于回答多项选择式的问题,表示提到的各项都成立。
So I think that all of your decisions when you're training a model come back to pre-training.
所以我认为,当你训练一个模型时,你所有的决策都会回到预训练上。
So if you're going to scale RL on a model, you still need to decide on your architecture that enables this.
所以如果你要在模型上扩展强化学习,你仍然需要决定能够实现这一点的架构。
We were talking about other architectures using different types of attention.
我们之前讨论过使用不同类型注意力的其他架构。
We're also talking about mixture of experts models.
我们也在讨论专家混合模型。
This sparse nature of MOE models makes it much more efficient to do generation,
MOE模型的这种稀疏特性使得生成更加高效,
which becomes a big part of post-training.
这成为后训练的重要组成部分。
And it's like, you need to have your architecture ready so that you can actually scale up this compute.
就像,你需要准备好你的架构,这样你才能真正扩展这个计算。
I still think most of the compute is going in at pre-training.
我仍然认为大部分计算都投入在预训练中。
Because you can still make a model better, you still want to go and revisit this.
因为你仍然可以让模型变得更好,你仍然想要去重新审视这一点。
You still want the best base model that you can.
你仍然想要你能得到的最好的基础模型。
And in a few years, that'll saturate and the RL compute will just go longer.
而在几年后,那会饱和,而强化学习的计算只会持续更久。
Is there people who disagree with you?
有人不同意你的观点吗?
that say basically pre-training is dead.
他们说基本上预训练已死。
It's all about scaling inference, scaling post-training, scaling context, continual learning, scaling data, synthetic data.
一切都关乎扩展推理、扩展后训练、扩展上下文、持续学习、扩展数据、合成数据。
People vibe that way and describe it in that way,
人们有那种感觉,并那样描述,
but I think it's not the practice that is happening.
但我认为实际发生的实践并非如此。
It's just the general vibe of people saying this thing is dead.
这只是人们说这个东西已死的普遍氛围。
The excitement is elsewhere.
兴奋点在其他地方。
So the low-hanging fruit in RL is elsewhere.
所以强化学习中的低垂果实在其他地方。
For example, we released our model in November for every company has deadlines.
例如,我们在11月发布了我们的模型,因为每家公司都有截止日期。
Our deadline was like November 20th.
我们的截止日期大约是11月20日。
And for that, our RL run was five days, which compared to 2024 is a very long time
为此,我们的强化学习运行了五天,与2024年相比,这已经是很长的时间了
to just be doing post-training at a model of like 30 billion parameters.
仅仅是在一个约300亿参数的模型上进行后训练。
It's not a big model.
这不是一个大模型。
And then in December, we had another release, which was just we let the RL run go for another three and a half weeks and the model got notably better.
然后在12月,我们又有一次发布,就是让强化学习运行再持续了三周半,模型明显变得更好了。
notably better 常用搭配
明显更好
用于描述某事物有显著提升,比 much better 更正式一点。
So we release it.
所以我们发布了它。
And that's a big amount of time to just allocate to something that is going to be your peak for the year.
而这是把大量时间仅仅分配给一个将成为你年度巅峰的东西。
allocate to 常用搭配
把(时间、资源)分配给
常用于讨论时间、资金或资源的分配,后接对象。
So it's like there's these types of discussions that happen when they're training a model where they just like can't they can't leave it forever
所以就像在训练模型时会出现这类讨论,他们就是不能永远放着不管
leave it forever 常用搭配
一直放着不管
口语中表示无限期搁置某事,常用于讨论不能拖延的情况。
you have to keep you have to keep pulling in the improvements you have from your researchers
你必须不断从研究人员那里获取改进
pulling in 常用搭配
不断获取、引入
口语中表示持续从某处获取资源或信息,如 pull in improvements。
so that's like you redo pre-training you'll do this post-training for a month but then you need to give it to your users you need to do safety testing
所以就像你重做预训练,你会做这个后训练一个月,但然后你需要把它交给用户,你需要做安全测试
so it was kind of just like i think there's a lot in place that reinforces this cycle of just keep updating the models there's things to improve you get a
所以这有点像,我认为有很多因素在强化这个循环,就是不断更新模型,总有东西可以改进,你会得到一个
there's a lot in place 常用搭配
有很多因素已经到位
用于说明某个系统或局面中已有许多支撑条件。
New compute cluster that lets you do something maybe more stable or faster.
新的计算集群,让你能做一些可能更稳定或更快的事情。
It's like you hear a lot about Blackwell having rollout issues where at AI2,
就像你听到很多关于Blackwell推出问题的消息,在AI2,
most of the models were pre-training around like one to two thousand GPUs.
大多数模型都是在约一到两千个GPU上进行预训练的。
But when you're pre-training on 10,000 or 100,000 GPUs, you hit very different failures.
但当你在1万或10万个GPU上进行预训练时,你会遇到非常不同的故障。
hit very different failures 常用搭配
遇到非常不同的故障
用于描述在不同规模或条件下出现不同类型的失败。
So GPUs are known to break in weird ways, and doing 100,000 GPU run is like you're pretty much guaranteed to always have at least one GPU that is down,
所以GPU以奇怪的方式出故障是出了名的,而进行10万GPU的运行,就像你几乎肯定总是至少有一个GPU是宕机的,
known to break in weird ways 常用搭配
以奇怪的方式出故障是出了名的
用于描述某物容易以意想不到的方式损坏,常带幽默或无奈语气。
and you need to have your training code handle that redundancy, which is just a very different problem.
而你需要让你的训练代码处理这种冗余,这只是一个非常不同的问题。
handle that redundancy 常用搭配
处理这种冗余
技术讨论中表示系统需要应对重复或备用机制。
Whereas like what we're doing, like I'm playing with post-training on DGX Spark,
而像我们正在做的,比如我在DGX Spark上玩后训练,
or you have your book, it's like, or people learning ML, it's like what they're battling to train these biggest models is just like mass distributed scale.
或者你有你的书,就像,或者人们学习机器学习,就像他们努力训练这些最大模型所面临的,就像大规模分布式扩展。
And it's a very different, but that's somewhat different than like are these like, that's a systems problem.
这是一个非常不同的,但那有点不同于像这些,那是一个系统问题。
In order to enable the scaling laws, especially a pre-training, you need all these GPUs at once.
为了实现扩展定律,尤其是预训练,你需要同时拥有所有这些GPU。
When we shift to reinforcement learning, it actually lends itself to heterogeneous compute,
当我们转向强化学习时,它实际上适合异构计算,
lends itself to 常用搭配
适合、适用于
表示某事物天然适合某种用途或方法。
because you have many copies of the model and to do a primer for a language
因为你有许多模型副本,并且为一个语言做入门
model reinforcement learning, what you're doing is you have two sets of GPUs.
模型强化学习,你要做的是拥有两组GPU。
One is you can call it the actor, and one you call the learner.
一组你可以称之为actor,另一组你称之为learner。
The learner is where your actual reinforcement learning updates are going to do these.
learner是实际强化学习更新将要进行的地方。
are traditionally policy gradient algorithms, proximal policy optimization, PPO, and group relative policy optimization, GRPO, are the two popular classes.
传统上是策略梯度算法,近端策略优化(PPO)和组相对策略优化(GRPO)是两种流行的类别。
And on the other side, you're going to have actors which are generating completions.
另一方面,你会有actor,它们正在生成补全。
And these completions are the things that you're going to grade.
这些补全就是你要评分的东西。
So reinforcement learning is all about optimizing reward.
所以强化学习就是关于优化奖励。
And in practice, what you can do is that you can have a lot of different actors in different parts of the world doing different types of problems.
在实践中,你可以做的是,你可以在世界不同地方有很多不同的actor,处理不同类型的问题。
And then you send it back to this highly networked compute cluster to do this actual learning.
然后你把它发送回这个高度网络化的计算集群来进行实际的学习。
where you take the gradients and you need to have a tightly meshed network where you can do different types of parallelism and spread out your model for efficient training.
在这里你获取梯度,你需要有一个紧密连接的网络,可以进行不同类型的并行计算,并分散你的模型以实现高效训练。
So there's just like a lot of every different type of training and serving has these considerations you need to scale.
所以就像每种不同类型的训练和服务都有这些你需要扩展的考虑因素。
Like we talked about pre-training, we talked about RL and then inference time scaling is like,
就像我们谈过的预训练,我们谈过强化学习,然后推理时间扩展就像是,
how do you serve a model that's thinking for an hour to 100 million users?
你如何为一个思考一小时的模型服务1亿用户?
I don't really know about that, but I know that's a hard problem.
我不太了解那个,但我知道那是个难题。
And in order to give people this intelligence, there's all the systems problems and we need more compute and you need more stable compute to do it.
而为了给人们这种智能,有所有的系统问题,我们需要更多计算,你需要更稳定的计算来做到。
you're bullish on all of these kinds of scaling is what i'm hearing on the inference on the reasoning even on the pre-training
你看好所有这些类型的扩展,这是我听到的,关于推理、关于推理能力,甚至关于预训练
yeah so that's a big can of worms here but so that
是的,所以那是个大麻烦,但所以那
a big can of worms 地道口语
一个大麻烦、复杂棘手的问题
口语中形容一旦处理就会引发许多麻烦的复杂问题。
basically two the knobs are the training and the inference scaling
基本上两个旋钮是训练和推理扩展
where you can get gains and so in a world where we had let's say infinite compute resources
在那里你可以获得收益,所以在一个我们拥有比如说无限计算资源的世界里
you want to do all of them about like so you have training you have inference scaling
你想做所有这些,就像所以你有训练,你有推理扩展
and training is like a hierarchy it's pre-training mid-training post-training changing the model size, more training data, training a bigger model.
而训练就像一个层级,它是预训练、中期训练、后训练,改变模型大小、更多训练数据、训练一个更大的模型。
gives you more knowledge in the model than the model
给模型带来的知识比模型
um let's say has a better, it's like a better base model back in the day, or we still we call it foundation model, and it unlocks
嗯,比如说有更好的,就像以前更好的基础模型,或者我们仍然称之为基础模型,而它解锁了
so you but you don't let's say have the model be able to solve your most complex task tasks during pre-training or after pre-training
所以,但假设你并没有让模型能够在预训练期间或预训练之后解决你最复杂的任务
you still have these other unlock phases where you have mid-training or long context for example post-training with lrvr that unlocks capabilities that the model has in terms of just knowledge in the pre-training
你仍然有这些其他的解锁阶段,比如你有中期训练或长上下文,例如使用 lrvr 的后训练,这解锁了模型在预训练中仅就知识而言所具备的能力
and I think sure if you so do more pre-training you get a better base model that you can unlock later
而且我认为,如果你做更多的预训练,你会得到一个更好的基础模型,之后你可以解锁它
but like Nathan said it just becomes too expensive so we don't have infinite compute so you have to decide do I want to spend that compute more on making the model larger
但就像 Nathan 说的,它变得太昂贵了,所以我们没有无限的计算资源,因此你必须决定:我是想把那些计算资源更多地花在让模型更大上
but you know it's like a trade-off, it's it's like in an ideal world you want to do all of them and I think in that sense scaling is still pretty much alive
但你知道,这就像一种权衡,就像在理想世界中你想做所有这些,而我认为从这个意义上说,扩展仍然相当活跃
a trade-off 常用搭配
一种权衡、取舍
用于说明在两种选择之间必须牺牲一方以换取另一方。
you would still get a better model but like we saw with gpd 4.5 it's just not worth it I mean it's like because you can let's
你仍然会得到一个更好的模型,但就像我们在 gpd 4.5 上看到的,它就是不值得,我的意思是,这就像因为你可以,让我们
not worth it 常用搭配
不值得
口语中表示做某事付出的代价大于收益。
Say you can unlock more performance with other techniques at that current moment,
比如说,你可以通过其他技术在当前时刻解锁更多性能,
especially, um, if you look at inference scaling,
尤其是,嗯,如果你看看推理扩展,
that's one of the biggest gains this year with 01, um,
这是今年最大的收益之一,用01,嗯,
where it took a smaller model further than pre-training a larger model like GBD 4.5.
它让一个较小的模型比预训练一个更大的模型(如GBD 4.5)走得更远。
So it's like I wouldn't say pre-training scaling is that it's just like there are other more attractive ways to scale right now at the moment,
所以这就像我不会说预训练扩展就是那样,只是现在有其他更有吸引力的扩展方式,
but at some point, you know, you will still want to make some progress on the pre-training.
但在某个时候,你知道,你仍然希望在预训练上取得一些进展。
The thing is also to consider, um, where you, where do you want to spend your money?
还要考虑的是,嗯,你,你想把钱花在哪里?
If you spend it more on the pre-training, it's like a fixed cost: you train the model and then it has this capability forever, you can always use it and so forth.
如果你在预训练上花更多钱,这就像固定成本:你训练模型,然后它就永远拥有这个能力,你可以一直使用它等等。
With inference scaling, you don't spend money during training, you spend money later per query, and then it's also like the math: how long is my model going to be on the market?
对于推理扩展,你不在训练期间花钱,而是之后按查询花钱,然后这还涉及数学:我的模型将在市场上存在多久?
If I replace it in half a year, maybe it's not worth spending five million, ten million, hundred million dollars on the training.
如果我半年内替换它,也许不值得在训练上花费五百万、一千万、一亿美元。
It longer, maybe it's just I will just do more inference scaling and get
它更长,也许我只是会做更多的推理扩展并获得
The performance from there, it may cost me two million in terms of user queries.
从那里的性能来看,就用户查询而言,它可能会花费我两百万。
It becomes a question of how many users you have and then doing the math.
这就变成了一个你有多少用户然后做数学计算的问题。
Um, and I think that's also where it's interesting.
嗯,我认为这也是有趣的地方。
Where GGPD is in a position, I think they have a lot of users.
GGPD 所处的位置,我认为他们有很多用户。
Where they need to go a bit cheaper.
他们需要变得更便宜一些。
Where they have that GPD5 model that is a bit smaller.
他们有那个稍微小一点的 GPD5 模型。
Other companies that have as if your customers have other trade-offs.
其他公司,如果你的客户有其他权衡的话。
For example, there was also the math Olympiad or some of these math problems.
例如,还有数学奥林匹克竞赛或一些这样的数学问题。
Where JGBT or OpenAI, they had a proprietary model.
JGBT 或 OpenAI,他们有一个专有模型。
And I'm pretty sure it's just like a model that has been maybe fine-tuned a little bit more,
而且我很确定它只是一个可能被稍微更多微调过的模型,
but most of it was during inference scaling to achieve this peak performance in certain tasks.
但大部分是在推理扩展期间实现的,以在某些任务中达到这种峰值性能。
Where you don't need that all the time.
你并不需要一直那样。
But yeah, long story short, I do think all of these pre-training, mid-training, post-training, infant scaling, they are all still things you want to do.
但是的,长话短说,我确实认为所有这些预训练、中期训练、后训练、婴儿扩展,它们仍然都是你想做的事情。
long story short 地道口语
长话短说
口语中用于总结前面冗长的内容,直接给出结论。
It's just finding at the moment in this year, it's finding the right ratio that gives you the best bang for the buck, basically.
只是今年此刻,基本上就是找到能给你带来最佳性价比的正确比例。
best bang for the buck 地道口语
性价比最高
口语中表示花最少的钱获得最大的回报。
I think this might be a good place to define pre-training, mid-training, and post-training.
我认为这可能是一个定义预训练、中期训练和后训练的好地方。
So pre-training is the classic training, one next token prediction at a time.
所以预训练就是经典的训练方式,一次预测下一个词元。
You have a big corpus of data.
你有一个庞大的数据语料库。
And Nathan also has very interesting insights there, because of OMO3, it's a big portion of the paper focuses on the right data mix.
而Nathan在这方面也有非常有趣的见解,因为OMO3,论文很大一部分都在关注正确的数据配比。
So pre-training is essentially just, you know, cross entropy loss training on next token prediction on a vast corpus of internet data, books, papers, and so forth.
所以预训练本质上就是,你知道,在庞大的互联网数据、书籍、论文等语料库上进行交叉熵损失训练,预测下一个词元。
It has changed a little bit over the years in the sense people used to throw in everything they can.
这些年来它有一点变化,就是过去人们会把所有能拿到的东西都扔进去。
Now, it's not just raw data.
现在,它不只是原始数据。
It's also synthetic data where people rephrase certain things.
它还包括合成数据,也就是人们对某些内容进行改写。
So synthetic data doesn't necessarily mean purely AI made up data.
所以合成数据并不一定意味着纯粹由AI编造的数据。
It's also taking something from an article, Wikipedia article, and then rephrasing it as a Q&A question or summarizing it, rewording it and making better data that way.
它也包括从一篇文章、维基百科文章中提取内容,然后把它改写成问答形式,或者进行总结、改写,从而制作出更好的数据。
Because I think of it also like with humans, if someone, let's say, reads the book compared to a messy, I don't know, no offense, but like Reddit post or something like that.
因为我也会把它类比到人类身上,如果一个人,比如说,读一本书,相比于一个杂乱的,我不知道,无意冒犯,但就像Reddit帖子之类的。
no offense 地道口语
无意冒犯
在说出可能让人不快的话之前先打个招呼,缓和语气。
I do think you learn, no offense, but I think.
我确实认为你能学到东西,无意冒犯,但我认为。
no offense 地道口语
无意冒犯
在说出可能让人不快的话之前先打个招呼,缓和语气。
There's going to be a post about this.
会有一篇关于这个的帖子。
Some Reddit data is very coveted and excellent for training.
有些Reddit数据非常抢手,非常适合训练。
You just have to filter it.
你只需要对它进行过滤。
I think that's the idea.
我觉得这就是那个思路。
I think it's like if someone took that and rephrases that in a, let's say, more concise and structured way,
我觉得就像如果有人把它拿来,用更简洁、更有结构的方式重新表述,
I think it's higher quality data that gets the LLM maybe the same.
我觉得那是更高质量的数据,让大语言模型也许得到同样的结果。
You get the same LLM out of it at the end, but it gets there faster.
你最终从中得到同样的大语言模型,但它到达得更快。
It trains faster because let's say if the grammar and the punctuation is correct,
它训练得更快,因为比如说如果语法和标点都是正确的,
it already learns the correct way versus getting information from a messy way and then learning later how to correct that and stuff like that.
它已经学到了正确的方式,而不是从混乱的方式中获取信息,然后再学习如何纠正它之类的。
and stuff like that 地道口语
以及诸如此类的东西
口语中列举完例子后收尾,表示还有其他类似的东西。
So I think that is how pre-training evolved and how still why scaling still works is that it's not about just the amount of data, it's also the tricks to make that data better for you in a sense.
所以我认为这就是预训练演变的方式,以及为什么扩展仍然有效,因为它不只是关于数据量,还在于让那些数据在某种意义上对你更有用的技巧。
And mid-training is, I mean, it used to be called pre-training.
而中期训练,我的意思是,它过去被称为预训练。
I think it's called mid-training because it was awkward to have pre-training and post-training, but nothing in the middle, right?
我认为它被称为中期训练,是因为有预训练和后训练,但中间什么都没有,这很别扭,对吧?
It sounds a bit weird. You have pre-training and post-training, but what's the actual training?
听起来有点奇怪。你有预训练和后训练,但实际的训练是什么?
So the mid-training is usually similar to pre-training, but you know, it's a bit more, I would say, specialized in pre-training.
所以中期训练通常和预训练类似,但你知道,它更,我会说,更专门化于预训练。
It's the same algorithm, but what you do is you focus, for example, on long context.
算法是一样的,但你要做的是专注于,比如说,长上下文。
Like one example, you have long context documents.
举个例子,你有长上下文文档。
The reason you don't do that during just pure pre-training is because you don't have that many long context documents.
你在纯预训练期间不这么做的原因是,你没有那么多长上下文文档。
So you have a specific phase.
所以你有一个特定的阶段。
And one problem of LLMs is also still, it's a neural network.
而LLMs的一个问题也仍然是,它是一个神经网络。
has the problem of catastrophic forgetting so you teach it something it forgets other things and you want to it's not 100 forgetting but you know it's like no free lunch you can't it's also the same with humans
有灾难性遗忘的问题,所以你教它一些东西,它会忘记其他东西,你想要——不是100%遗忘,但你知道,就像没有免费的午餐,你不能——人类也是一样
no free lunch 常用搭配
没有免费的午餐,凡事都有代价
表示得到某样东西必然要付出代价或有所牺牲。
if you ask me some math i learned 10 years ago i don't know i would have to look at it again
如果你问我一些我10年前学的数学,我不知道,我得再看一遍
nathan was actually saying that he's consuming so much content that there's a catastrophic forgetting issue
内森实际上在说,他消费了太多内容,以至于出现了灾难性遗忘的问题
yeah i'm like trying to learn so much about ai i was like i was learning about pre-training parallelism i'm like i lost something and i don't know what it was
是的,我就像试图学习很多关于AI的东西,我当时正在学习预训练并行性,我就像我丢失了一些东西,我不知道那是什么
i don't anthropomorphize llms but it's i think the same kind of in that sense how humans learn i mean
我不会将LLMs拟人化,但我认为在那种意义上,这就像人类学习一样,我是说
the quantity is not always better because yeah you it's like being selective and i in the mid training is being selective in terms of quality content at the end
数量并不总是更好,因为是的,你——这就像要有选择性,而我在中期训练中就是在质量内容方面要有选择性,在最后
so the last thing the lm has seen is the quality stuff and then post training is all
所以LM最后看到的东西是高质量的内容,然后后训练全是
the uh fine-tuning, supervised fine-tuning, DPO, um, reinforcement learning with verifiable rewards, with human feedback, and so forth.
呃,微调、监督微调、DPO、嗯,带可验证奖励的强化学习、带人类反馈的强化学习等等。
So the refinement stages, and it's also interesting, it's like the cost thing, right? I mean,
所以这些精炼阶段,而且这也很有意思,就像成本问题,对吧?我是说,
it's like pre-training, you spend a lot of money on that, right?
就像预训练,你在那上面花很多钱,对吧?
Now RL, a bit less. RL, you don't really...
现在强化学习,稍微少一点。强化学习,你其实不……
I would say teach it knowledge. It's more like unlocking the knowledge.
我会说是教它知识。这更像是解锁知识。
It's more like a skill learning, like how to solve problems with the knowledge that it has from pre-training.
这更像是一种技能学习,比如如何用它在预训练中获得的知识来解决问题。
There are actually three papers this year or last year, 2025, on RL for pre-training.
实际上今年或去年,2025年,有三篇关于用强化学习做预训练的论文。
But I, I mean, I don't think anyone does that in production, toy examples for now.
但,我是说,我觉得没有人在生产环境里那么做,目前只是玩具示例。
Examples, right? But to generalize, RL post-training is more like the skill unlock,
示例,对吧?但要泛化的话,强化学习后训练更像是解锁技能,
where pre-training is like soaking up the knowledge, essentially.
而预训练本质上就像是吸收知识。
soaking up the knowledge 常用搭配
大量吸收知识
比喻像海绵一样不断吸收、积累知识。
A few things that could be helpful for people, a lot of people get like...
有几件事可能对人们有帮助,很多人会像……
You have, think of synthetic data as being bad for training the models, you mentioned, like
你有,把合成数据看作对训练模型有害,你提到过,就像
the DeepSeek get almost OCR, which is optical character recognition paper.
DeepSeek 那篇几乎就是 OCR,也就是光学字符识别的论文。
A lot of labs did. AI2 had one, had multiple. And the reason that each of these labs have these is because
很多实验室都做过。AI2 有一个,有多个。而这些实验室之所以都有这些,是因为
There's vast amounts of PDFs and other digital documents on the web
网上有大量的PDF和其他数字文档
that are in formats that aren't encoded with text easily, so you use these almost OCR, these or deep-seek OCR
它们的格式不容易用文本编码,所以你要用这些几乎算是OCR的,这些或者DeepSeek OCR
and we called our almost OCR to extract
我们把自己的这个几乎算是OCR的东西叫做提取
what can be trillions of tokens of candidate data for pre-training
可以提取出数万亿个token的候选预训练数据
and pre-training data set size is on the order of trillions, is measured in trillions of tokens
而预训练数据集的大小是数万亿这个量级,是以数万亿个token来衡量的
Smaller models from researchers can be something like five to ten trillion, um, quen is documented
研究人员做的小模型可能是五到十万亿,嗯,Qwen是有记录的
going up to like 50 trillion, and there's rumors that these closed labs can go to like 100 trillion tokens
能到大概50万亿,还有传言说这些闭源实验室能做到大概100万亿个token
and just getting this potential data to put in, I think they, they have a very big funnel
而只是获取这些潜在数据放进去,我觉得他们,他们有一个非常大的漏斗
and then the data you actually train the model on is a small percentage of this, like the, this character recognition data
然后你实际用来训练模型的数据只是其中的一小部分,就像这个字符识别数据
would be described as synthetic data for pre-training in a lab, and then
在实验室里会被描述为用于预训练的合成数据,然后
there's also the things like ChatGPT now
还有像现在ChatGPT这样的东西
gives wonderful answers and you can train on those best answers and that's synthetic data
能给出很棒的答案,你可以用那些最好的答案来训练,那就是合成数据
it's very different than like early ChatGPT, lots of hallucinations data
这和早期的ChatGPT很不一样,那时候有很多幻觉数据
when people became grounded synthetic data.
当人们变得有依据时的合成数据。
One interesting question is, if I recall correctly, Olmo 3 was trained with less data than specifically some other open-weight models, maybe even Olmo 2,
一个有趣的问题是,如果我没记错的话,Olmo 3 的训练数据比某些其他开放权重模型要少,甚至可能比 Olmo 2 还少,
if I recall correctly 地道口语
如果我没记错的话
在陈述记忆中的信息时,表示不确定、留有余地。
but you still got better performance and that might be one of the examples how the data helped.
但你还是得到了更好的表现,这也许就是数据如何发挥作用的一个例子。
It's mostly down to data quality.
这主要取决于数据质量。
I think if we had more compute, we would train for longer.
我想如果我们有更多算力,我们会训练更长时间。
I think we ultimately see that as, like, just like something we would want to do
我想我们最终会把这看作,就像,我们想做的一件事,
and especially with big models, you need to have more compute
尤其是对于大模型,你需要有更多算力,
because we talk about having more parameters and we talk about knowledge and essentially there's a ratio where big models can absorb more from data
因为我们谈论拥有更多参数,我们谈论知识,本质上存在一个比例,大模型可以从数据中吸收更多,
and then you're going to you get more benefit out of this
然后你就能从中获得更多收益,
it's it's like one of these any logarithmic graph in your mind is like a small model
这就像你脑海中的那种对数图,小模型
will level off sooner if you're measuring trends of tokens and bigger and bigger models
如果你衡量的是 token 的趋势,会更早趋于平缓,而越来越大的模型
need more but mostly is we aren't training that big of models right now ai2
需要更多,但主要是我们现在并没有训练那么大的模型,AI2
and getting the highest quality data we can is the natural starting point is there
而尽可能获取最高质量的数据是自然的起点,有没有
something to be said uh about the topic of data quality is there some long-hanging fruit there still where the quality could be improved
关于数据质量这个话题还有什么可说的,是否还有一些容易摘取的果实,质量还能提高
long-hanging fruit 常用搭配
容易摘取的果实,容易实现的目标
指最容易达成、投入最少就能获得成果的部分。
it's like turning the crank
这就像转动曲柄一样
So I think historically in the open, there's been like a canonical
所以我认为从历史上看,在开放领域,一直有一个像是经典的
best pre-training data set that has moved around between who has the most recent one or the best recent effort.
最好的预训练数据集,它在谁拥有最新的或最好的近期成果之间流转。
Like AI2's Dolmo was very early with the first Ulmo and Hugging Face had FineWeb.
比如AI2的Dolmo很早就推出了第一个Ulmo,而Hugging Face有FineWeb。
And there's a DCLM project, which has been kind of like a, which is, it stands for data comp language model.
还有一个DCLM项目,它有点像是,也就是,它代表数据竞赛语言模型。
There's been data comp for other machine learning projects and they have had a very strong data set.
其他机器学习项目也有数据竞赛,它们拥有非常强大的数据集。
and a lot of it is the internet is becoming fairly closed off
而且其中很多是互联网正变得相当封闭
so we have common crawl which i think is hundreds of trillions of tokens and you filter it and it looks like being
所以我们有Common Crawl,我认为它有数百万亿个token,你过滤它,它看起来像是
a lot of scientific work where you're training classifiers and making decisions based on how do you prune down this this data set into the highest quality stuff and the stuff that suits your tasks
很多科学工作,你训练分类器并基于如何将这个数据集修剪成最高质量的内容以及适合你任务的内容来做决定
prune down 常用搭配
精简、削减(数据或内容)
用于描述从大量内容中删减筛选,保留最有价值的部分。
so previously language models were tested a lot more on like knowledge and just kind of conversational things but now they're expected to do math and code.
所以以前语言模型更多是在知识和对话类内容上测试,但现在它们被期望做数学和代码。
were tested a lot more on 句型
过去更多是在……方面被测试
[subject] were tested a lot more on [something]
用于对比过去和现在对某事物的测试或评估重点。
So to train a reasoning model, you need to remix your whole data set.
所以训练一个推理模型,你需要重新混合你的整个数据集。
And there's a lot of actually wonderful scientific methods here where you can take your gigantic data set, you sample a lot of really tiny things from different sources.
这里有很多实际上很棒的科学方法,你可以取用你庞大的数据集,从不同来源采样很多非常小的东西。
sample a lot of really tiny things from different sources 常用搭配
从不同来源抽取大量很小的样本
用于描述从多个来源少量取样以构建数据集的做法。
So you say you have GitHub, Stack Exchange, Reddit, Wikipedia,
所以你说你有 GitHub、Stack Exchange、Reddit、维基百科,
you can sample small things from them and train small models on each of these mixes
你可以从中抽取小样本,并在每个混合数据上训练小模型,
and measure their performance on your evaluations.
然后在你的评估集上衡量它们的表现。
And you can just do like basic linear regression and it's like, here's your optimal data set.
你可以直接做基本的线性回归,然后它就会说,这是你的最优数据集。
But if your evaluations change, your data set changes a lot.
但如果你的评估变了,你的数据集就会变化很大。
So a lot of Olmode 3 was new sources for reasoning to be better at math and code.
所以 Olmode 3 的很多内容都是新的推理来源,以便在数学和代码方面表现更好。
And then you do this mixing procedure and it gives you the answer.
然后你执行这个混合过程,它就会给你答案。
And I think that's a lot of that's happened at labs this year.
我认为今年实验室里发生了很多这样的事。
There's new hot things, whether it's coding environments or web navigation.
有新的热门方向,无论是编码环境还是网页导航。
new hot things 地道口语
新的热门事物或方向
口语中用来指当下很受关注、很流行的事物。
You just need to bring in new data.
你只需要引入新数据。
bring in new data 常用搭配
引入新数据
用于描述为项目或模型补充新的数据来源。
You need to change your whole pre-training so your post-training can work better and stuff like this.
你需要改变整个预训练,这样你的后训练才能更好地工作,诸如此类。
And that's like the constant re-evolution and the redetermining of what they care about for their models.
这就像是不断的重新演化,以及重新确定他们关心模型的哪些方面。
Are there fun anecdotes of what sources of data are particularly high quality that we wouldn't expect?
有没有一些有趣的轶事,关于哪些数据来源质量特别高,而我们原本不会想到?
You mentioned Reddit sometimes can be a source.
你提到 Reddit 有时可以成为一个来源。
Reddit was very useful.
Reddit 非常有用。
I think that PDFs is definitely one.
我认为 PDF 绝对是一个。
Especially archive.
尤其是 archive。
Yeah, so like AI2 has run Semantic Scholar for a long time, which is a, like you can say, is a competitor to Google Scholar with a lot more features.
是的,所以像 AI2 已经运行 Semantic Scholar 很长时间了,它可以说是一个 Google Scholar 的竞争对手,而且功能多得多。
And to do this, AI2 has found and scraped a lot of PDFs
为了做到这一点,AI2找到并抓取了很多PDF文件
for openly accessible papers that might not be like behind the closed paid garden of a certain publisher.
这些是开放获取的论文,可能不在某个出版商封闭的付费花园后面。
behind the closed paid garden 常用搭配
在封闭的付费围墙之后
比喻内容被付费墙或封闭平台限制,外人无法自由获取。
So like truly open scientific PDFs.
所以就像是真正开放的科学PDF。
And if you like you sit on all of these and you process it and you can get value out of it.
如果你坐在所有这些数据上,处理它们,你就能从中获得价值。
get value out of it 常用搭配
从中获得价值
用于说明通过处理某物能得到有用的收益或成果。
And I think that like a lot of that style of work has been done by the frontier labs much earlier.
我认为很多这种风格的工作早就被前沿实验室做过了。
And it's just like you need to have a pretty skilled researcher that understands how things change models and they bring it in and they clean it.
这就像你需要一个相当熟练的研究员,了解模型如何变化,他们把它带进来并清理它。
And that's a lot of labor that like I think of a lot of frontier labs when they scale researchers a lot more goes into data.
这是大量的劳动,就像我认为很多前沿实验室在扩大研究人员规模时,更多的精力投入到数据中。
You have people like if you want to if you join a frontier lab, you want to have impact.
你有这样的人,如果你想加入一个前沿实验室,你想要产生影响。
The best way to do it is just make find new data that's better.
最好的方法就是找到更好的新数据。
And then like the fancy, glamorous algorithmic things like figuring out how to make a one is like the sexiest thought of a scientist.
然后像那些花哨、迷人的算法之类的东西,比如弄清楚如何做一个,是科学家最性感的想法。
It's like, oh, I figured out to scale RL.
就像,哦,我弄明白了如何扩展强化学习。
And there's a group that did that.
有一个团队做到了这一点。
But I think most of the contributions is like I'm going to make the data better or I'm going to make the infrastructure better
但我认为大部分的贡献就像是,我要把数据做得更好,或者我要把基础设施做得更好,
so that everybody in my team can run experiments five percent faster.
这样我团队里的每个人都能把实验跑快百分之五。
At the same time, I think it's also one of the closest guarded secrets what your training data is for legal reasons.
与此同时,我认为它也是最严守的秘密之一,出于法律原因,你的训练数据是什么。
closest guarded secrets 常用搭配
最严守的秘密
用于形容被极力保护、不对外公开的信息。
And so there's also, I think, a lot of work that goes into hiding what your training data was, essentially.
所以我认为,也有很多工作本质上是为了隐藏你的训练数据是什么。
Like trying the model to not give away the sources because of legal reasons.
就像试图让模型不要泄露来源,因为法律原因。
give away 常用搭配
泄露、透露
用于表示无意或有意地暴露本应保密的信息。
The other thing to be complete is that some people are trying to train on only licensed data,
另一件需要补充完整的事情是,有些人正试图只用获得许可的数据来训练,
where Common Crawl is a scrape of like the whole internet.
而Common Crawl基本上是对整个互联网的抓取。
So if I host multiple websites, I'm happy to have them train language models, but I'm not explicitly licensing what governs it.
所以如果我托管多个网站,我愿意让它们训练语言模型,但我并没有明确地授权什么来约束它。
And therefore, the common crawl is largely unlicensed, which means that your consent really hasn't been provided for how to use the data.
因此,Common Crawl基本上是没有许可的,这意味着关于如何使用这些数据,其实并没有得到你的同意。
There's another idea where you can train language models only on data that has been licensed explicitly.
还有另一种想法,就是你可以只用明确获得许可的数据来训练语言模型。
So that kind of governing contract is provided.
这样就提供了一种约束性的合同。
And I'm not sure if Aperitice is the copyright thing or the license thing.
我不确定Aperitice是版权方面的事还是许可方面的事。
I know that the reason that they did it was for an EU compliance thing
我知道他们这么做的原因是为了一个欧盟合规的事情
where they wanted to make sure that their model fit one of those checks.
他们想确保他们的模型符合其中一项检查。
And also on that note, also, for example, there's also the distinction between the licensing.
另外关于这一点,还有,例如,还有许可之间的区别。
So some people, like you said, they just purchase the license.
所以有些人,就像你说的,他们只是购买许可证。
Let's say they buy a book online, let's say an Amazon Kindle book or let's say a money book or something, and then use that in the training data.
假设他们在网上买一本书,比如亚马逊Kindle电子书,或者比如一本关于钱的书之类的,然后把它用在训练数据里。
And that is like the gray zone because you paid for the content and you might want to train it.
这就像是灰色地带,因为你为内容付了钱,你可能想用它来训练。
gray zone 常用搭配
灰色地带
指法律或规则上界定不清、存在争议的模糊区域。
But then there are also restrictions where even that shouldn't be allowed.
但还有限制,甚至那也不应该被允许。
And so that is like where it gets a bit fuzzy.
所以这就是它变得有点模糊的地方。
gets a bit fuzzy 地道口语
变得有点模糊不清
口语中用于描述情况或界限变得不明确、难以界定。
And yeah, I think that is right now still a hot topic.
是的,我认为现在这仍然是一个热门话题。
hot topic 常用搭配
热门话题
指当前被广泛讨论、备受关注的话题。
And also big companies like OpenAI, they approached private companies for their proprietary data.
还有像OpenAI这样的大公司,他们联系私营公司获取他们的专有数据。
And private companies, they become more and more, let's say, protective of their data because they know, okay, this is going to be my mode in a few years.
而私营公司,他们变得越来越,比如说,保护自己的数据,因为他们知道,好吧,这将在几年内成为我的模式。
And I do think that's like the interesting question:
我确实认为那才是有趣的问题:
where if LLMs become more commoditized,
如果大语言模型变得更加商品化,
and I think a lot of people learn about LLMs there,
而且我认为很多人会在那里了解大语言模型,
there will be a lot more people able to train LMs, of course.
当然,会有更多人能够训练语言模型。
There are infrastructure challenges, but if you think of big industries like pharmaceutical industries, law, finance industries,
存在基础设施方面的挑战,但如果你想想制药、法律、金融这样的大行业,
I do think they at some point will hire people from other frontier labs
我确实认为他们到某个时候会从其他前沿实验室招人,
to build their in-house models on their proprietary data, which will be then again another unlock with pre-training
用他们的专有数据构建自己的内部模型,而这又会成为预训练带来的又一次突破,
that is currently not there, because even if you wanted to, you can't get that data, you can't get access to clinical trials most of the time in these types of things.
而这种突破目前还不存在,因为即使你想做,你也拿不到那些数据,在这类事情中大多数时候你无法获得临床试验的数据。
So I do think scaling in that sense might be still pretty much alive,
所以我确实认为,从这个意义上说,规模化可能仍然相当有生命力,
if you also look in domain-specific applications,
如果你也看看特定领域的应用,
because we are still right now in this year just looking at general-purpose LMs on ChatGPT, Anthropic and so forth.
因为我们今年现在仍然只是在看ChatGPT、Anthropic等上的通用语言模型。
They are just general purpose, they're not even, I think, scratching the surface of what an LM can do if it is really specifically trained and designed for a specific task.
它们只是通用的,我认为它们甚至还没有触及语言模型在真正针对特定任务专门训练和设计时所能做到的事情的表面。
scratching the surface 常用搭配
触及表面,仅了解皮毛
用于表示只接触到事物的一小部分,远未深入。
I think on the data thing, this is one of the things where this happened in 2025 and we totally forget it, is Anthropic lost in court and was owed $1.5 billion to authors.
我认为在数据这件事上,这是2025年发生而我们完全忘记的事情之一,就是Anthropic在法庭上败诉,被判向作者支付15亿美元。
Anthropic, I think, bought thousands of books and scanned them and was cleared legally for that,
Anthropic,我认为,买了成千上万本书并扫描了它们,而且这在法律上是被允许的,
because they bought the books and that is kind of going through the system.
因为他们买了这些书,这算是走正规流程。
And then the other side, they also torrented some books.
而另一方面,他们还盗版下载了一些书。
And I think this torrenting was the path where the court said that they were then culpable to pay this billions of dollars to authors,
我认为这种盗版下载正是法院判定他们需要向作者支付数十亿美元的路径,
which is just like such a mind-boggling lawsuit that kind of just came and went.
这真是一场令人难以置信的诉讼,就这么来了又走了。
came and went 常用搭配
来了又走了,转瞬即逝
用于描述某事发生后又很快过去,未引起持续关注。
Like that is so much money from the VC ecosystem.
那从风投生态系统中可是很大一笔钱。
These are core cases that will define the future of human civilization,
这些是核心案例,将定义人类文明的未来,
because it's clearly that data drives a lot of this.
因为很明显,数据驱动了其中很多方面。
And there's this very complicated human tension of,
而且存在这种非常复杂的人性张力,
I mean, you can empathize, you're both authors.
我的意思是,你可以感同身受,你们俩都是作者。
There's some degree to which, I mean, you put your heart and soul and your sweat and tears into the writing that you do.
在某种程度上,我的意思是,你把自己的心血、灵魂、汗水和泪水都倾注到你写的作品里。
put your heart and soul 常用搭配
倾注全部心血和感情
形容对某件事投入极大的热情和精力,常用于写作、创作等语境。
It feels a little bit like theft for somebody to train your data without giving you credit.
别人用你的数据训练却不给你署名,这感觉有点像盗窃。
giving you credit 常用搭配
给予你认可或署名
指在使用他人作品或数据时标明来源、承认贡献,常用于学术或创作语境。
And like Nathan said, also two layers to it.
就像Nathan说的,这也有两个层面。
Someone might buy the book and then train on it, which could be argued fair or not fair.
有人可能会买这本书然后用来训练,这可以说公平,也可以说不公平。
could be argued 句型
可以说是有争议的
could be argued [fair or not fair]
用于表示某件事存在不同看法,可被支持或反对。
But then there are literally straight up companies who use pirated books where it's not even compensating the author.
但然后确实有一些公司直接使用盗版书籍,甚至都不补偿作者。
straight up 地道口语
直接地、毫不掩饰地
口语中用来强调某事的直接或明显程度,常带批评语气。
That is, I think, where people got a bit angry about it specifically.
我认为,这正是人们对此特别生气的地方。
Yeah, but there has to be some kind of compensation scheme.
是的,但必须有某种补偿机制。
This is like moving towards something like Spotify streaming did originally for music.
这就像朝着Spotify最初为音乐所做的流媒体模式发展。
What does that competition look like?
那种竞争是什么样的?
You have to define those kinds of models.
你必须定义那些类型的模型。
You have to think through all of that.
你必须把这一切都考虑清楚。
think through 常用搭配
仔细考虑、全面思考
指对复杂问题进行全面、深入的思考,常用于需要周密计划的语境。
One other thing I think people are generally curious about, I'd love to get your thoughts.
我觉得人们普遍好奇的另一件事,我很想听听你的想法。
I'd love to get your thoughts 地道口语
我很想听听你的想法
礼貌地征求对方意见,常用于访谈或讨论中。
As LLMs are used more and more, if you look at even Archive, but GitHub, more and more of the data is generated by LLMs.
随着LLM被越来越多地使用,如果你看看甚至Archive,还有GitHub,越来越多的数据是由LLM生成的。
What do you do in that kind of world?
在那种世界里你该怎么办?
How big of a problem is that?
这个问题有多大?
largest problems infrastructure and systems but from an ai point of view it's kind of inevitable
最大的问题在于基础设施和系统,但从AI的角度来看,这有点不可避免
so it's basically llm generated data that's curated by humans essentially right yes
所以基本上就是由人类策划的LLM生成数据,对吧,是的
and i think that a lot of open source contributors are legitimately burning out
而且我认为很多开源贡献者确实正在精疲力竭
burning out 常用搭配
精疲力竭、耗尽精力
指因长期压力或过度工作而身心疲惫,常用于工作或志愿贡献语境。
if you have a popular open source repo
如果你有一个受欢迎的开源仓库
somebody's like oh i want to do open source ai it's good for my career and
有人会说哦我想做开源AI,这对我的职业有好处,然后
they just vibe code something and they throw it into the
他们只是凭感觉写点代码,然后把它扔进
vibe code 地道口语
凭感觉写代码,不严谨地编程
非正式说法,指不经过深思熟虑、靠直觉快速写代码。
you might get more of this than i do
你可能比我得到更多这种东西
So I have actually a case study here.
所以我这里其实有一个案例研究。
I have a repository called ML Extend that I developed as a student, I don't know, 15 years, 10 years ago.
我有一个叫ML Extend的仓库,是我当学生时开发的,我不知道,15年、10年前吧。
And it's a reasonably popular library still for certain algorithms, I think, especially like frequent data mining stuff.
它仍然是一个相当受欢迎的库,用于某些算法,我想,尤其是像频繁数据挖掘之类的东西。
And there was recently, I think, two or three people who submitted a lot of PRs in a very short amount of time.
最近,我想,有两三个人在很短的时间内提交了很多PR。
I do think LMS have been involved in submitting these PRs.
我确实认为LMS参与了提交这些PR。
Me as the maintainer, there are two things.
作为维护者,我有两点。
First, I'm a bit overwhelmed.
首先,我有点不知所措。
a bit overwhelmed 常用搭配
有点不知所措、应付不过来
形容因事情太多或压力太大而感到难以应对。
I don't have time to read through it because especially it's an older library that is not a priority for me.
我没有时间通读,尤其是因为它是一个较旧的库,对我来说不是优先事项。
At the same time, I kind of also appreciate it because I think something people forget is it's not just using the LLM.
同时,我也有点感激,因为我认为人们忘记的是,这不仅仅是使用LLM。
There's still a human, you have a human layer that verifies something.
仍然有人,你有一个验证某些东西的人类层。
And that is in a sense also how data is labeled, right?
从某种意义上说,这也是数据被标注的方式,对吧?
So that's like one of the most expensive things is getting labeled data for RL back in human feedback phases.
所以这就像最昂贵的事情之一,就是在人类反馈阶段为RL获取标注数据。
And this is kind of like that where it goes through phases and then you get actually higher quality data out of it.
这有点像那样,它经历多个阶段,然后你从中得到更高质量的数据。
So I don't mind it in a sense.
所以从某种意义上说,我不介意。
It can feel overwhelming, but I do think there is also value in that.
它可能让人感到不知所措,但我确实认为其中也有价值。
It feels like there's a fundamental difference between raw LLM-generated data and LLM-generated data with human in the loop that does some kind of verification, even if that verification is a small percent of the lines of code.
感觉原始LLM生成的数据和有人类参与循环进行某种验证的LLM生成数据之间存在根本区别,即使这种验证只占代码行的一小部分。
human in the loop 常用搭配
人在回路中,即有人工参与验证
技术语境中指在自动化流程中保留人工审核或干预环节。
I think this goes with anything where people think also sometimes,
我认为这适用于任何人们有时也会认为的情况,
oh, I can just use an LLM to learn about XYZ, which is true, you can,
哦,我可以只用LLM来学习XYZ,这是真的,你可以,
but there might be a person who is an expert who might have used an LLM to write specific code.
但可能有一个专家,他可能用LLM写了特定的代码。
There is kind of like this human work that went into it to make it nice and throwing out the not so nice part to make it to kind of like pre-digest it for you.
有一种人类的工作投入其中,使其变得好,并扔掉不太好的部分,使其为你预先消化。
And that saves you time.
这节省了你的时间。
And I think that's the value add where you have someone filtering things or even using the LLMs correctly.
我认为这就是增值之处,有人过滤信息,甚至正确使用LLM。
value add 常用搭配
增值、附加价值
指某事物带来的额外价值或改进,常用于商业或产品语境。
I think this is still labor that you get for free with you, for example, read an article, let's say a sub-stack article,
我认为这仍然是你免费获得的劳动,例如,读一篇文章,比如说一篇sub-stack文章,
I could maybe ask an LLM to give me opinions on that, but I wouldn't even maybe know what to ask.
我可能可以让LLM给我关于那篇文章的意见,但我甚至可能不知道问什么。
I think there is still value in reading that article compared to me going to the LLM, because you are the expert, you select what knowledge is actually spot on, should be included, and you give me this executive summary.
我认为读那篇文章仍然有价值,相比于我去LLM,因为你是专家,你选择哪些知识实际上是准确的、应该包括的,然后你给我这个执行摘要。
spot on 地道口语
完全正确、非常准确
口语中表示某事物完全正确或恰到好处。
And this is kind of a huge value add, because now I don't have to waste three, five hours to go through this myself.
这是一个巨大的增值,因为现在我不必浪费三、五个小时自己通读这个。
maybe get some incorrect information and so on and so
也许得到一些错误的信息等等,所以
I think that's also where the future still is for writers
我认为这也正是作家未来的所在
even though there are LLMs that expert can kind of like save your time
即使有些大语言模型能帮你节省时间
it's kind of fascinating to actually watch and
实际观察起来还挺有意思的,而且
I'm sure you guys do this but for me to look at the difference between the summary and the original content
我相信你们也会这么做,但对我来说,去看摘要和原始内容之间的差异
even if it's a page-long summary of a page long content.
即使它是一页长内容的摘要,也有一页那么长。
It's interesting to see how the summary LMB summary takes the edge off.
有意思的是,看看摘要,LMB摘要,是如何削弱锋芒的。
takes the edge off 常用搭配
削弱锋芒、缓和强度
指使某事物不那么强烈、尖锐或令人不适。
Like what, what is the signal it removes from the thing?
比如,它从内容中移除了什么信号?
The voice is what I talk about a lot. Voice.
声音就是我经常谈论的东西。声音。
Well, voice, I would love to hear what you mean by voice.
嗯,声音,我很想听听你说的声音是什么意思。
That's really powerful, but sometimes there's like literally insights, like in removing an insight,
这确实很有力,但有时候真的会有一些洞见,比如在移除一个洞见时,
you're actually fundamentally changing the meaning of the thing.
你实际上从根本上改变了这个东西的含义。
So I'm continuously disappointed how bad LM's are at really getting to the core insights,
所以我一直很失望,大语言模型在真正抓住核心洞见方面有多差,
which is what a great summary does.
而这正是好的摘要该做的事。
Even if you go, and I have these extensive, extremely elaborate prompts where I'm like really trying to dig for the insights
即使你走了,而我有这些详尽、极其精细的提示,我真的很努力去挖掘那些见解
and it's still not quite there, which, I mean,
但它仍然没有完全达到,我的意思是,
that's a whole deep philosophical question about what is human knowledge and wisdom and what does it mean to be insightful and so on.
那是一个完整的深刻哲学问题,关于什么是人类的知识和智慧,以及富有洞察力意味着什么等等。
But when you talk about the voice, what do you mean?
但当你谈到声音时,你是什么意思?
So when I write, I think a lot of what I'm trying to do is take what you think as a researcher,
所以当我写作时,我认为我试图做的很多事就是把你作为研究者的想法,
which is very raw, which a researcher is trying to encapsulate an idea at the frontier of their understanding.
那是非常原始的,研究者试图在他们理解的前沿概括一个想法。
And they're trying to put what is a feeling into words.
他们试图把一种感觉用语言表达出来。
put what is a feeling into words 句型
把一种感觉用语言表达出来
put [something] into words
用于描述将抽象感受转化为具体表达的困难过程。
And I think that my writing, I tried to do this as the writing, which makes it come across as raw,
我认为我的写作,我试图这样做,这使得它显得原始,
but also high information in a way that it's like some people will get it and some won't.
但也信息量很高,就像有些人会理解,有些人不会。
And that's kind of the nature of research.
这有点像是研究的本质。
And I think this is something that language models don't do well.
我认为这是语言模型做得不好的地方。
Particularly, they're all trained with this reinforcement learning from human feedback, which is designed to take feedback from a lot of people and, in a way, average how the model behaves from this.
特别是,它们都通过这种从人类反馈中进行强化学习来训练,这种学习旨在从很多人那里获取反馈,并在某种程度上平均模型的行为。
And I think that it's going to be hard for a model to be very incisive when there's that sort of filter in it.
我认为当模型中有那种过滤时,它很难变得非常犀利。
And I think this is kind of a wonderful fundamental problem for researchers in RLHF.
我认为这对RLHF的研究者来说是一个很美妙的基础性问题。
It's like this provides so much utility in making the models better.
就像这为让模型变得更好提供了很多效用。
But also the problem formulation is kind of like there's this knot in it that you can't get past.
但问题的表述方式也有点像里面有个你无法解开的结。
So that's what I think of is like these language models don't have this prior and their deep expression that they're trying to get at.
所以我认为这些语言模型没有这种先验,也没有它们试图表达的深层表达。
I don't think it's impossible to do.
我不认为这是不可能做到的。
I think there's stories of models that really shock people.
我认为有些模型的故事真的让人震惊。
I think of like I would love to have tried Bing Sydney.
我想,比如我很想试试Bing Sydney。
And does that have more voice?
那个有更多声音吗?
Because it would so often go off the rails on people in what is historically obviously a scary way, like telling a reporter to leave its wife is a crazy model to potentially put in general adoption.
因为它经常会在人们身上失控,以一种历史上明显可怕的方式,比如告诉记者离开它的妻子,这是一个可能投入普遍使用的疯狂模型。
go off the rails 常用搭配
失控、偏离正轨
形容人或系统行为突然变得异常、不受控制,常用于口语。
But that's kind of like a trade-off, like is this RLHF process, like in some ways adding limitations.
但这有点像一种权衡,比如这个RLHF过程,在某些方面增加了限制。
a trade-off 常用搭配
权衡、取舍
指为了得到某样东西而不得不放弃另一样,常用于讨论利弊。
That's a terrifying place to be as one of these frontier labs and companies because millions of people are using them.
作为这些前沿实验室和公司之一,这是一个可怕的位置,因为数百万人在使用它们。
There was a lot of backlash last year with the GPT-4-0 getting removed.
去年有很多反对声音,因为GPT-4-0被移除了。
a lot of backlash 常用搭配
大量反对、强烈反弹
指公众对某个决定或变化表达强烈不满,常用于新闻或讨论中。
And I personally never used the model,
而我个人从未使用过这个模型,
but I've talked to people at OpenAI
但我与OpenAI的人交谈过,
where they're to the point
他们到了这样的地步,
where they get emails from users
用户会给他们发邮件,
that might be detecting subtle differences in the deployments in the middle of the night.
这些邮件可能是在深夜检测到部署中的细微差异。
And they email them and they're like, my friend is different.
他们给他们发邮件,说,我的朋友不一样了。
And they find these employees' emails and send them things
他们找到这些员工的邮箱,给他们发东西,
because they're so attached to this set of model weights and a configuration that is deployed to the users.
因为他们如此依恋这套部署给用户的模型权重和配置。
attached to 常用搭配
对……依恋、离不开
形容对某人或某物产生情感依赖,常用于口语。
We see this with TikTok.
我们在TikTok上看到了这一点。
You open it. I don't use TikTok.
你打开它。我不用TikTok。
Supposedly in like five minutes, the algorithm gets you.
据说大约五分钟,算法就能抓住你。
the algorithm gets you 地道口语
算法就把你抓住了
口语中形容推荐算法很快让人上瘾、被吸引住。
It's locked in.
它就被锁定了。
locked in 常用搭配
被锁定、被牢牢抓住
形容无法摆脱某种状态,常用于口语。
And I don't, like, those are language models doing recommendations.
而我不,就像,那些是语言模型在做推荐。
Like, I think there are ways that you can do this with a language model.
就像,我认为有办法用语言模型做到这一点。
Within, like, five minutes of chatting with it, the model just gets you.
在,就像,与它聊天五分钟内,模型就能抓住你。
the model just gets you 地道口语
模型就是懂你
口语中形容某人或某物非常理解你,这里指AI模型能抓住你的心思。
And that is something that people aren't really ready for.
而这是人们还没有真正准备好的事情。
Like, I think, like, don't give that to kids.
就像,我认为,就像,不要把这个给孩子。
Like, don't give that to kids at least until we know what's happening.
就像,至少在我们知道发生了什么之前,不要把这个给孩子。
But there's also going to be this mechanism.
但也会有这种机制。
What's going to happen with these LLMs is they're used more and more.
这些LLM将会发生的是,它们被越来越多地使用。
Unfortunately, the nature of the human condition is such that people commit suicide.
不幸的是,人类境况的本质就是人们会自杀。
And so what journalists would do is they will report extensively on the people who commit suicide.
所以记者们会做的是,他们会广泛报道那些自杀的人。
And they would very likely link it to the LLMs because they have that data about the conversations.
他们很可能会把它和LLM联系起来,因为他们有那些对话数据。
If you're really struggling in your life, if you're depressed, if you're thinking about suicide, you're going to probably talk to LLMs about it.
如果你在生活中真的苦苦挣扎,如果你抑郁,如果你在考虑自杀,你很可能会和LLM谈论这些。
And so what journalists will do is they will say, well, the suicide was committed because of the LLM.
所以记者们会做的是,他们会说,嗯,自杀是因为LLM。
And that's going to lead to the companies, because of legal issues and so on, more and more and more taking the edge off of the LLM.
而这将导致公司,因为法律问题等等,越来越削弱LLM的锋芒。
taking the edge off 常用搭配
削弱锋芒、降低强度
指让某事物不那么尖锐或强烈,常用于口语。
So it's going to be as generic as possible.
所以它会尽可能通用。
It's so difficult to operate in this space because, of course, you don't want an LLM to cause harm to humans at that level.
在这个领域运作非常困难,因为当然,你不希望LLM对人类造成那种程度的伤害。
But also, this is also the nature of the human experience, is to have a rich conversation, a fulfilling conversation, one that challenges you from which you grow.
但同样,这也是人类体验的本质,就是进行丰富的对话、充实的对话,一个挑战你并让你成长的对话。
You need that edge.
你需要那种锋芒。
And that's something extremely difficult for AI researchers on the RLHF front to actually have to solve.
而这是AI研究人员在RLHF方面必须解决的极其困难的事情。
Because you're actually dealing with the human condition.
因为你实际上是在处理人类境况。
Like a lot of researchers at these companies are so well motivated.
就像这些公司的很多研究人员都非常有动力。
And there's definitely the likes of Anthropic and OpenAI are culturally so want to do good through this for the world.
而且像Anthropic和OpenAI这样的公司,在文化上确实非常想通过这个为世界做好事。
And it's such a, I'm like, I don't want to work on this.
而且这真是,我就像,我不想做这个。
Because on the one hand, a lot of people see AI as a health ally, as somebody they can talk to about their health confidentially.
因为一方面,很多人把AI视为健康盟友,视为可以私下谈论自己健康的人。
But then it bleeds all the way into this talking about mental health and things where it's heartbreaking that this will be the thing where somebody goes over the edge.
但接着它又蔓延到谈论心理健康之类的事情上,令人心碎的是,这将成为有人崩溃的导火索。
goes over the edge 常用搭配
崩溃、越过临界点
形容人情绪或行为失控,常用于心理健康语境。
But other people might be saved.
但其他人可能会因此得救。
And I'm like, I don't.
而我会想,我不想这样。
There's things that as a researcher training models, it's like I don't want to train image generation models and release them openly.
有些事情,作为一名训练模型的研究员,就像我不想训练图像生成模型并公开释放它们。
Because I don't want to enable somebody to have a tool on their laptop that can harm other people.
因为我不想让某人拥有一个能伤害他人的笔记本电脑工具。
Like, I don't have the infrastructure at my company to do that safely.
就像,我的公司没有基础设施来安全地做到这一点。
But it's like, there's a lot of areas like this where it's just, it needs people that will approach it with the complexity and just kind of conviction of like, it's just such a hard problem.
但就像,有很多这样的领域,它只是需要人们以复杂性和那种信念来对待它,就像,这只是一个非常困难的问题。
But also we as a society, as users of these technologies, need to make sure that we're having the complicated conversation about it versus just fear-mongering.
但同样,我们作为社会,作为这些技术的用户,需要确保我们进行的是关于它的复杂对话,而不是仅仅散布恐惧。
fear-mongering 常用搭配
散布恐惧、制造恐慌
指通过夸大危险来引发恐惧,常用于批评性讨论。
Big tech is causing harm to humans or stealing your data, all that kind of stuff. It's more complicated than that.
大型科技公司正在对人类造成伤害或窃取你的数据,诸如此类。事情比这更复杂。
And you're right.
你说得对。
There's a very large number of people inside these companies,
这些公司内部有非常多的人,
many of which you know, many of which I know, that deeply care about helping people.
其中许多你认识,许多我也认识,他们深切关心帮助他人。
They are considering the full human experience of people from across the world, not just Silicon Valley.
他们正在考虑来自世界各地人们的完整人类体验,而不仅仅是硅谷。
People across the United States, people across the world, what that means, what their needs are.
美国各地的人们,世界各地的人们,这意味着什么,他们的需求是什么。
It's really difficult to design this one system
设计这样一个系统真的很难,
that is able to help all these different kinds of people across the different age groups, cultures, mental states, mental conditions, all that kind of stuff.
它能够帮助所有这些不同种类的人,跨越不同的年龄组、文化、心理状态、精神状况,所有这类东西。
I wish that the timing of AI was different with the relationship of big tech to the average person.
我希望人工智能的时机有所不同,考虑到大型科技公司与普通人之间的关系。
So like big tech's reputation was so low.
所以就像大型科技公司的声誉如此之低。
And with how AI is so expensive, it's like inevitably going to be a big tech thing where it takes so many resources.
而且由于人工智能如此昂贵,它不可避免地会成为大型科技公司的事情,需要消耗如此多的资源。
And people say the U.S. is quote unquote betting the economy on AI with this build out.
人们说美国所谓地在用这次建设将经济押注在人工智能上。
quote unquote 地道口语
所谓的、带引号的
口语中用于表示某个词是别人说的或带有讽刺意味。
And it's like to have these be intertwined at the same time is just makes for such a hard communication environment.
而让这些同时交织在一起,只会造成如此艰难的沟通环境。
It would be good for me to go talk to more people in the world that hate big tech and see AI as a continuation of this.
对我来说,去和世界上更多憎恨大型科技公司并将人工智能视为其延续的人交谈会很好。
And one of the things you actually recommend, one of the antidotes that you talk about.
而你实际推荐的一件事,你谈到的一种解药。
is to find agency in this whole system,
是在整个系统中找到能动性,
as opposed to sort of sitting back in a powerless way
而不是以一种无能为力的方式坐视不管,
and consuming the AI slop as it quickly, rapidly takes over the internet.
并消费那些AI垃圾,因为它迅速、快速地接管了互联网。
More find agency by using it to build stuff, build apps, build.
更多地通过用它来构建东西、构建应用、构建来找到能动性。
So one, that actually helps you build intuition, but two, it's empowering
所以第一,这实际上帮助你建立直觉,但第二,它赋予你力量,
because you can understand how it works, what the weaknesses are, and
因为你可以理解它是如何工作的,弱点是什么,并且
it gives your voice power to say, like, this is fucked up, this is bad, this is bad use of the technology and this is good use of technology
它赋予你的声音力量去说,比如,这太糟糕了,这很糟糕,这是对技术的糟糕使用,而这是对技术的良好使用,
and you're more plugged into the system then so you can understand it better
而且你更深入地接入这个系统,因此你能更好地理解它,
and you can steer it better as i think it's a good point you brought up agency
并且你能更好地引导它,我认为你提出的能动性这一点很好,
instead of ignoring it and saying okay i'm not going to use it i think it's probably long-term healthier
而不是忽视它并说好吧我不会使用它,我认为长期来看更健康的是,
to say okay it's out there i can't put it back you know like internet computers back then when they came out
说好吧它已经存在了我无法收回它,你知道就像互联网电脑刚出现时那样,
how do i make best use of it and how does it help me to up level myself
我如何最好地利用它,它如何帮助我提升自己,
up level myself 常用搭配
提升自己
口语中表示提高自己的能力或水平。
the one thing i worry here though is like if you just fully use it for something you love to do the the thing you love to do is not no longer there and
不过这里我担心的一件事是,如果你完全用它来做你热爱的事情,你热爱的事情就不再存在了,而且
That could potentially, I feel like, lead to burnout, for example.
我觉得那可能会导致倦怠,比如说。
lead to burnout 常用搭配
导致倦怠
指长期压力或过度工作引发身心疲惫,常用于职场或健康讨论。
If I use an LLM to do all my coding for me,
如果我用大语言模型替我完成所有编码,
now there's no coding, I'm just managing something that is coding for me, two years, let's say, later.
现在没有编码了,我只是在管理一个替我编码的东西,比如说两年后。
If I just do that eight hours a day, I have something code for me, do I feel
如果我每天就那样做八小时,有东西替我编码,我会觉得
fulfilled? Still like, is this like, yeah, I mean, is this like
满足吗?还是像,这像,对,我是说,这像
hurting me in terms of being excited about my job, excited about what I'm doing, am I still proud to build something?
在让我对工作感到兴奋、对我所做的事感到兴奋方面伤害我吗?我还为构建某个东西感到自豪吗?
in terms of 常用搭配
在……方面;就……而言
用于引出讨论的某个方面或角度,口语和书面都常用。
So there's a, on that topic of enjoyment, it's quite interesting.
所以关于享受这个话题,有个,挺有意思的。
We should just throw this in there. There's this recent survey of about 791 professional developers, professional meaning 10 plus years of experience. That's a long time, yeah.
我们应该把这个放进去。最近有一项调查,大约791名专业开发者,专业指的是10年以上经验。那是很长的时间,对。
throw this in there 地道口语
把这个也加进去/顺便提一下
口语中表示在讨论或内容中顺便加入某个话题或信息。
As a junior developer, uh, yeah, in this day and age, uh, so the, the results here on many fronts are, uh, surprising.
作为一名初级开发者,呃,对,在这个时代,呃,所以这里在许多方面的结果,呃,令人惊讶。
in this day and age 常用搭配
在当今这个时代
用于强调当前时代的特点,常带感慨或对比意味。
So they break it down by junior and senior developers. But I mean, it just shows that both junior, senior developers use AI-generated code in code they ship. So this is not just for fun, sort of intermediate kind of learning things. This is code they ship.
所以他们按初级和高级开发者进行了细分。但我的意思是,这只是表明初级和高级开发者都在他们发布的代码中使用AI生成的代码。所以这不仅仅是为了好玩,那种中级的、学习性质的东西。这是他们要发布的代码。
break it down 常用搭配
把它分解/细分
表示把复杂信息按类别或部分拆开分析,常用于解释数据或问题。
And so it's 25%. Most of them use around 50% or more.
所以是25%。他们大多数人使用大约50%或更多。
And what's interesting is for the category of over 50% of your code that you ship is AI-generated,
有趣的是,对于你发布的代码中超过50%是AI生成的这个类别来说,
senior developers are much more likely to do so.
高级开发者更有可能这样做。
But you don't want AI to take away the thing you love.
但你不想让AI夺走你热爱的东西。
I think this speaks to my experience, these particular results I'm about to say.
我觉得这印证了我的经历,也就是我接下来要说的这些具体结果。
speaks to my experience 常用搭配
印证了我的经历
表示某个结果、说法与自己的亲身经历相符。
So together, about 80% of people find it either somewhat more enjoyable or significantly more enjoyable to use AI as part of the work?
所以总的来说,大约80%的人觉得把AI作为工作的一部分来使用要么稍微更愉快,要么明显更愉快?
I think it depends on the task.
我觉得这取决于任务。
For my personal usage, for example, I have a website where I sometimes tweak things on the website.
就我个人的使用来说,比如,我有一个网站,我有时会在网站上调整一些东西。
I personally don't enjoy this. So in that sense,
我个人并不喜欢这个。所以从这个意义上说,
in that sense 常用搭配
从这个意义上说
用于承接前文,从某个特定角度得出结论或限定说法。
if the AI can help me to implement something on my website,
如果AI能帮我在我的网站上实现一些东西,
I'm all here for it. It's great.
我完全支持。这很棒。
I'm all here for it 地道口语
我完全支持/非常乐意
口语中表示对某事物十分赞同或欢迎。
But then at the same time, when I solve a complex problem,
但与此同时,当我解决一个复杂问题时,
well, if there's a bug and I hunt this bug and I find the bug, it's the best feeling in the world.
嗯,如果有一个bug,我追踪这个bug并找到了它,那是世界上最好的感觉。
It's like you get so much joy, like oh, it's like you feel like great.
就像你获得了如此多的快乐,就像哦,就像你感觉棒极了。
But now if you don't even think about thinking about the bug, you just go directly to the LLM,
但现在如果你甚至不考虑思考这个bug,你直接去找LLM,
well, you never have this kind of feeling, right?
嗯,你永远不会有这种感觉,对吧?
But then there could be the middle ground where well, you try yourself, you can't find it, you use the LLM,
但然后可能有一个中间地带,嗯,你自己尝试,找不到它,你使用LLM,
the middle ground 常用搭配
中间地带;折中方案
指两种极端之间的折中做法或立场。
and then you don't get frustrated because it helps you and you move on to something that you enjoy and so
然后你不会感到沮丧,因为它帮助你,你继续做你喜欢的事情,所以
So I think looking at these statistics, I think also the difference is what is not factored in.
所以我认为看这些统计数据,我认为区别还在于没有考虑到什么。
what is not factored in 常用搭配
没有被考虑进去的因素
用于指出分析或统计中遗漏、未纳入考量的因素。
It's averaging over all the different scenarios where we don't know if it's for the core task or if it's for something mundane that people would not have enjoyed otherwise.
它是对所有不同场景的平均,我们不知道它是用于核心任务,还是用于人们本来不会喜欢的平凡事情。
So in a sense, AI is really great for doing mundane things that take a lot of work.
所以从某种意义上说,AI真的非常擅长做那些需要大量工作的平凡事情。
So, for example, my wife the other day, she has like a podcast for like book discussions, a book club.
所以,例如,前几天我妻子,她有一个像书籍讨论的播客,一个读书俱乐部。
And she was like transferring show notes from Spotify to YouTube.
她当时正在把节目笔记从Spotify转移到YouTube。
And then the links somehow broke.
然后那些链接不知怎么的就失效了。
And she had in some episodes, because it is customary books, like 100 links or something.
而且她在某些剧集里,因为按惯例书籍之类的,大概有100个链接。
And it would have been really painful to go in there and fix each link manually.
要进去手动修复每个链接会非常痛苦。
And so I suggested, hey, let's try ChatGPT.
所以我建议,嘿,我们试试ChatGPT吧。
We copied the text into ChatGPT and it fixed them.
我们把文本复制到ChatGPT里,它就把它们修好了。
And instead of two hours going from link to link, fixing that, you know, it made that type of work much more seamless.
而不是花两个小时从一个链接到另一个链接地修复,你知道,它让这类工作变得顺畅多了。
There was no frustration fixed.
没有挫败感,修好了。
I think everyone has a use case where AI is useful for something like that, that would be really boring, really mundane.
我认为每个人都有这样的用例,AI对类似的事情很有用,那些事情会非常无聊、非常平凡。
use case 常用搭配
使用场景;用例
指某项技术或产品适合被使用的具体情境。
I, for me personally, since we're talking about coding, and you mentioned debugging,
我,就我个人而言,既然我们在谈论编程,而且你提到了调试,
a lot of the source of the enjoyment for me, more on the cursor side than the cloud code side,
对我来说,很多乐趣的来源,更多是在cursor这边而不是cloud code那边,
is the, I have a friend, I have a co, what's that called, a pair programmer.
是,我有一个朋友,我有一个合,那叫什么来着,结对程序员。
It's less lonely.
这样没那么孤独。
You made debugging sound like this great joy.
你把调试说得像是一种巨大的乐趣。
No, I would say debugging is like a drink of water after you've been going through a desert for days.
不,我会说调试就像在沙漠里走了几天后喝到一口水。
So you skip the whole desert part where you're suffering.
所以你跳过了整个你受苦的沙漠部分。
So sometimes it's nice to have a friend who can't really find the bug, but can give you some intuition about the code.
所以有时候有个朋友很好,他可能找不到bug,但能给你一些关于代码的直觉。
And you're together with that friend going through the desert.
你和那个朋友一起穿越沙漠。
and then together find that drink of water.
然后一起找到那口水。
So at least for me, maybe it speaks to the loneliness of the programming experience.
所以至少对我来说,也许它道出了编程体验中的孤独感。
That is a source of joy.
那是一种快乐的来源。
It's maybe also related to delayed gratification.
它也许还和延迟满足有关。
delayed gratification 常用搭配
延迟满足
指为了更大的回报而忍受等待或克制即时欲望。
I'm a person who, you know, even as a kid, I like the idea of Christmas presents,
我是一个人,你知道,即使小时候,我也喜欢圣诞礼物的那种感觉,
having them, getting them better than actually getting the presents.
拥有它们、得到它们,比真正拿到礼物更好。
I would look forward to the day I get the presents, but then it's over and I'm disappointed.
我会期待收到礼物的那一天,但之后一切就结束了,我会感到失望。
look forward to 常用搭配
期待;盼望
表示对将来要发生的事感到期待,后接名词或动名词。
And maybe it's something like also with, let's say food.
也许这也类似于,比如说食物。
I think food tastes better when you're really hungry.
我觉得当你真的很饿的时候,食物尝起来更好吃。
And with, yeah, you're right with debugging.
还有,是的,你说得对,调试也是这样。
It is not always, you know, great.
它并不总是,你知道,很棒。
It's like often frustrating.
它常常令人沮丧。
But then if you can solve it, then it's great.
但如果你能解决它,那就很棒。
But there's also like a sweet Goldilocks zone if it's too hard and it's, you know, wasting your time.
但也有一个甜蜜的黄金地带,如果它太难,你知道,就是在浪费时间。
Goldilocks zone 常用搭配
恰到好处的理想区间
借童话《金发姑娘》比喻不太难也不太易、刚刚好的状态。
But I think that is another challenge though.
但我觉得那是另一个挑战。
How will people learn?
人们将如何学习?
I mean, the chart we looked at, we saw that more senior developers are shipping more AI-generated code than the junior ones.
我的意思是,我们看过的那个图表显示,更多资深开发者比初级开发者发布更多AI生成的代码。
And I think it's very interesting because intuitively you would think it's the junior developers
我觉得这很有趣,因为直觉上你会认为是初级开发者,
because they don't know, let's say, how to do the thing yet because they are more junior.
因为他们还不知道,比如说,怎么做那件事,因为他们更初级。
And so they use AI to do that thing.
所以他们用AI来做那件事。
It could either mean the AI is not good enough yet to solve that task,
这可能意味着AI还不够好,无法解决那个任务,
but it could also mean experts are more effective at using it.
但也可能意味着专家更擅长使用它。
They know where and better how to use it and review the code and they trust the code then more.
他们知道在哪里以及如何更好地使用它,审查代码,然后更信任代码。
And so I think one issue in the society in the future will be, though,
所以我认为未来社会的一个问题是,
how do you become an expert if you never try to do the thing yourself?
如果你从不自己尝试做那件事,你怎么成为专家?
And I think one way it's always like for me, how I learn is by trying things myself, like math textbooks.
我认为对我来说,一直以来的一个方式,我学习的方式就是自己尝试,比如数学教科书。
If you look at the solutions, yeah, you learn something.
如果你看解答,是的,你学到一些东西。
But I think you learn actually better if you try first and then you appreciate the solution differently
但我认为如果你先尝试,然后你会以不同的方式欣赏解答,实际上学得更好,
because you know how to put it into your mental framework.
因为你知道如何把它放入你的思维框架中。
And if LLMs are here all the time, would you actually go through the length at struggling?
如果LLM一直存在,你真的会经历挣扎的过程吗?
Would you be willing to struggle?
你愿意挣扎吗?
Because struggle is not nice, right? I mean, it's struggling.
因为挣扎并不好受,对吧?我是说,那就是挣扎。
I mean 地道口语
我是说;我的意思是
用于口语中补充、澄清或强调前面的话,使语气更自然。
And if you use the LLM to do everything at some point, you will never really take the next step.
如果你在某个时候用LLM做所有事情,你永远不会真正迈出下一步。
at some point 常用搭配
在某个时候
表示不确定的某个时间点,常用于谈论未来可能发生的事。
take the next step 常用搭配
迈出下一步
用于描述在过程或发展中进入下一个阶段。
And then you will maybe not get that unlock that you would get as an expert using an LLM.
然后你可能就不会获得那种解锁,那种你作为专家使用LLM时会获得的解锁。
So it's like, you know, it's like, I think there's like a Goldilocks sweet spot where maybe the trick here is you make dedicated offline time
所以就像,你知道,就像,我觉得有一个像金发姑娘的甜蜜点,也许这里的诀窍是你专门留出离线时间
sweet spot 常用搭配
最佳平衡点;理想状态
指各方面条件恰到好处、效果最好的中间状态。
the trick here is 句型
这里的诀窍是
the trick here is [to do something / that ...]
用于引出解决问题或达成目标的关键做法。
where you study two hours, a day and the rest of the day use llms
你每天学习两小时,其余时间使用LLM
but i think it's important also for people to still invest in themselves in my opinion
但我认为,在我看来,人们仍然投资自己也很重要
invest in themselves 常用搭配
投资自己
指花时间、精力或金钱提升自身能力或福祉。
in my opinion 地道口语
在我看来;依我看
用于表达个人观点,语气较礼貌、缓和。
to not just you know llm everything
不要只是你知道的什么都交给LLM
yeah there is and uh we together civilization that we each individually have to find that godilog zone
是的,有,而且呃我们共同文明,我们每个人都必须找到那个godilog区域
yeah uh and in the program and context as developers now we've had this fascinating conversation that started with pre training and mid training
是的呃,在程序和上下文中,作为开发者,我们现在进行了这个迷人的对话,从预训练和中训练开始
let's get to post training a lot of fun stuff in post training
让我们进入后训练,后训练中有很多有趣的东西
So what are some of the interesting ideas in post-training?
那么后训练中有哪些有趣的想法?
The biggest one from 2025 is learning this reinforcement learning with verifiable rewards.
2025年最大的一个是学习这种带有可验证奖励的强化学习。
You can scale up the training there, which means doing a lot of this kind of iterative generate grade loop.
你可以在那里扩大训练规模,这意味着进行大量这种迭代生成评分循环。
scale up 常用搭配
扩大规模;提升规模
用于描述增加资源、数量或范围以提升产出或能力。
And that lets the models learn both interesting behaviors on the tool use and software side.
这让模型在工具使用和软件方面都能学到有趣的行为。
This could be searching, running commands on their own and seeing the outputs.
这可能是搜索、自行运行命令并查看输出。
And then also that training enables this inference time scaling very nicely.
然后,这种训练也非常好地实现了推理时间扩展。
And it just turned out that this paradigm was very nicely linked in this, where it's this kind of RL training enables inference time scaling.
而事实证明,这种范式与此非常完美地联系在一起,即这种RL训练实现了推理时间扩展。
turned out 常用搭配
结果是;事实证明
用于说明最终发现或实际发生的情况与预期可能不同。
But inference time scaling could have been found in different ways.
但推理时间扩展本可以通过不同的方式发现。
So it was kind of this perfect storm of the models change a lot in the way that they're trained is a major factor in doing so.
所以这有点像一场完美风暴:模型在训练方式上发生了很大变化,这是实现这一目标的主要因素。
perfect storm 常用搭配
完美风暴;多种因素叠加导致的局面
指多个因素同时作用,造成特别显著或严重的结果,可褒可贬。
And this has changed how people approach post-training dramatically.
这极大地改变了人们进行后训练的方式。
Can you describe RLVR, popularized by DeepSeq R1?
你能描述一下由DeepSeq R1推广的RLVR吗?
Can you describe how it works?
你能描述一下它是如何工作的吗?
Yeah, fun fact, I was on the team that came up with the term RLVR,
是的,有趣的是,我是在提出RLVR这个术语的团队中,
fun fact 地道口语
有趣的是;说个有趣的事
用于引出一个有趣但可能不太重要的信息,常见于口语。
which is from our two to three work before DeepSeek,
这来自我们在DeepSeek之前的两到三篇工作,
which is we don't take a lot of credit for the being the people to popularize the scaling RL.
那就是我们并不因成为推广扩展RL的人而居功。
take a lot of credit for 常用搭配
因……而居功;把功劳归于自己
常用于否定句,表示不把某事的功劳算在自己头上。
But as fun as what academics get as an aside is the ability to name and influence the discourse
但和学术界作为副业所获得的一样有趣的是,命名和影响话语的能力
as an aside 常用搭配
顺便说一句;作为题外话
用于引入与主题相关但不直接的内容。
because the closed labs can only say so much that one of the things you can do as an academic is like you might not have the compute to train the model,
因为封闭实验室能说的有限,作为学者你能做的一件事就是,你可能没有算力来训练模型,
but you can frame things in a way that ends up being I describe it as like a community can come together around this RLVR term, which is very fun.
但你可以以某种方式构建事物,最终就像我描述的那样,一个社区可以围绕这个RLVR术语聚集起来,这非常有趣。
come together around 常用搭配
围绕……聚集起来;团结在……周围
指一群人因某个想法、目标或术语而形成共识或合作。
And then DeepSeq is the people that did the training breakthrough, which is they scaled the reinforcement learning,
然后DeepSeq是那些实现了训练突破的人,也就是他们扩展了强化学习,
which was you have the model generate answers and then grade the completion if it was right.
也就是让模型生成答案,然后对完成情况进行评分,看是否正确。
And then that accuracy is your reward for reinforcement learning.
然后这个准确率就是强化学习的奖励。
So reinforcement learning is classically an agent that acts in an environment and the environment gives it a state and a reward back.
所以强化学习经典上是一个在环境中行动的智能体,环境给它一个状态和一个奖励作为反馈。
And you try to maximize this reward.
然后你试图最大化这个奖励。
In the case of language models, the reward is normally accuracy on a set of verifiable tasks, whether it's math problems, coding tasks.
在语言模型的情况下,奖励通常是在一组可验证任务上的准确率,无论是数学问题还是编码任务。
And it starts to get blurry with things like factual domains,
然后它开始变得模糊,比如事实性领域,
get blurry 常用搭配
变得模糊;界限不清
用于描述概念、界限或区分不再清晰明确。
like that is also in some ways verifiable or constraints on your instruction,
比如那在某些方面也是可验证的,或者对你的指令有约束,
like respond only with words that start with A.
比如只用以A开头的词来回应。
Like all of these things are verifiable in some way.
就像所有这些事情在某种程度上都是可验证的。
And the core idea of this is you find a lot more of these problems that are verifiable
而它的核心思想是,你找到更多可验证的问题,
and you let the model try it many times while taking these RL steps,
然后你让模型尝试很多次,同时采取这些强化学习步骤,
these RL gradient updates, the infrastructure evolved from this reinforced learning from human feedback,
这些强化学习梯度更新,基础设施是从这种基于人类反馈的强化学习演变而来的,
where in that era, the score they were trying to optimize was a learned reward model of aggregate human preferences.
在那个时代,他们试图优化的分数是一个学习到的人类偏好聚合奖励模型。
So you kind of change the problem domains and that let the optimization go on to much bigger scales,
所以你有点改变了问题领域,这让优化能够扩展到更大的规模,
which kind of kickstarted a major change in what the models can do and how people use them.
这某种程度上开启了模型能做什么以及人们如何使用它们的重大变化。
What kind of domains is RLVR amenable to?
RLVR适合哪些领域?
Math and code are the famous ones.
数学和代码是著名的例子。
And then there's a lot of work kind of on what is called the rubrics,
然后有很多工作是关于所谓的评分标准,
which is related to a word people might have heard as L, I'm as a judge,
这与人们可能听过的“L,我作为评委”这个词有关,
which is like for each problem, I'll have a set of problems in my training data set.
这就像对于每个问题,我的训练数据集中会有一组问题。
I'll then have another language model and ask it, what would a good answer to this problem look like?
然后我会用另一个语言模型,问它:这个问题的好答案应该是什么样?
And then you can try the problem a bunch of times over and over again and assign a score based on this rubric.
然后你可以反复多次尝试这个问题,并根据这个评分标准打分。
over and over again 常用搭配
反复地;一次又一次
强调重复多次做某事。
So that's not necessarily verifiable, like a math and code domain.
所以这不一定像数学和代码领域那样可验证。
But this rubrics idea and other scientific problems that it might be a little bit more vague is where a lot of the attention is, where they're trying to push this set of methods into these kind of more open-ended domains where the models can learn a lot more.
但这种评分标准思路以及其他可能更模糊的科学问题,正是很多关注所在,他们试图把这套方法推向这些更开放式的领域,让模型能学到更多。
I think that's called reinforcement running with AI feedback, right?
我想那叫做带AI反馈的强化学习,对吧?
That's the older term from it that was coined in Anthropics constitutional AI paper.
那是更早的术语,出自Anthropic的宪法AI论文。
So it's like a lot of these things come in cycles.
所以就像很多这些东西都是循环出现的。
come in cycles 常用搭配
循环出现;周期性出现
指事物按周期反复出现,而非一次性发生。
just one step back for the RLVR.
只是为RLVR退一步。
So I think the interesting, beautiful thing here is that you ask the LLM, let's say, a math question, and then you know the correct answer.
所以我认为这里有趣而美妙的是,你问LLM,比如说,一个数学问题,然后你知道正确答案。
And you let the LLM, like you said, figure it out. But how it does it, I mean, you don't really constrain it much.
你让LLM,就像你说的,自己解决。但它是怎么做到的,我的意思是,你并没有真正限制它太多。
figure it out 常用搭配
弄明白;想出解决办法
指通过思考或尝试自己找到答案或方法。
There are some constraints you can add, like use the same language, don't switch between Spanish and English.
你可以加一些限制,比如使用同一种语言,不要在西班牙语和英语之间切换。
But let's say you're pretty much hands off, you only give the question and the answer, and then the LM has to, you know, just the task to arrive at the right answer.
但假设你基本上放手不管,你只给出问题和答案,然后语言模型必须,你知道,只是完成任务来得出正确答案。
hands off 常用搭配
放手不管;不干预
指不介入或尽量少干预,让对方自行处理。
But the beautiful thing here is what happens in practice is that the LM will do a step-by-step description, like you know, like as a student or like as a mathematician, how you would derive the solution.
但这里美妙的地方在于,实践中发生的是,语言模型会做一个逐步的描述,就像你知道的,像一个学生或者像一个数学家那样,你会如何推导出解决方案。
step-by-step 常用搭配
一步一步的;逐步的
用于描述按顺序、分步骤进行的过程或说明。
It will give you, it will use those steps and that helps actually the model to improve its own accuracy.
它会给你,它会使用那些步骤,而这实际上帮助模型提高自身的准确性。
And then, like you said, the inference scaling, so inference scaling loosely means basically spending more compute during using the LM during inference.
然后,就像你说的,推理扩展,所以推理扩展大致上意味着基本上在使用语言模型进行推理期间花费更多的计算资源。
And here the inference scaling is that the model would use more tokens.
而这里的推理扩展就是模型会使用更多的词元。
And also, I think in our one paper, they showed the longer they train the model, the longer the responses are, they grow over time, they use more tokens, so it becomes more expensive, becomes more expensive for simple tasks.
而且,我认为在我们的一篇论文中,他们展示了训练模型的时间越长,回复就越长,它们会随着时间增长,使用更多的词元,所以变得更昂贵,对于简单任务来说变得更昂贵。
But these explanations, they help the model with the accuracy.
但这些解释,它们帮助模型提高准确性。
There are also interesting, a lot of papers showing what the model explains does not necessarily have to be correct, or maybe it's even unrelated to the answer.
也有很多有趣的论文表明,模型解释的内容不一定正确,或者甚至可能与答案无关。
But for some reason, it still helps the model, like the fact that it is explaining.
但出于某种原因,它仍然对模型有帮助,比如它在解释这一事实。
And I think it's also, again, I don't want to anthropomorphize these LLMs, but it's kind of like how we humans operate, right?
而且我认为,再次强调,我不想把这些大语言模型拟人化,但这有点像我们人类运作的方式,对吧?
kind of like 地道口语
有点像;差不多像
口语中用于缓和语气,表示大致相似或不太确定的类比。
If there's a complex math problem, let's say in a math class, you usually have a notepaper and you do it step by step.
如果有一个复杂的数学问题,比如说在数学课上,你通常会有一张草稿纸,然后一步一步地做。
You cross out things. And the model also self-corrects.
你划掉一些东西。而模型也会自我纠正。
cross out 常用搭配
划掉;删去
指用线划掉文字或内容,表示取消或修改。
And that was, I think, the aha moment in the R1 paper.
而我认为,那就是R1论文中的顿悟时刻。
They called it aha moment because the model itself recognized it, made a mistake, and then said, ah, I did something wrong, and so let me try.
他们称之为顿悟时刻,因为模型自己意识到它犯了一个错误,然后说,啊,我做错了什么,所以让我试试。
I think that's just so cool that this falls out of just giving it the correct answer and having it figure out how to do it, that it kind of does, in a sense, what a human would do.
我觉得这太酷了,仅仅给它正确答案并让它自己找出如何做,就能产生这种效果,从某种意义上说,它有点像人类会做的事。
falls out of 常用搭配
由……自然产生;从……中得出
指某结果作为自然或附带的结果出现。
in a sense 常用搭配
在某种意义上
用于表示从某个角度看或某种程度上成立。
Although LMS don't think like humans. It's kind of like an interesting coincidence.
尽管语言模型不像人类那样思考。这有点像是一个有趣的巧合。
And the other nice side effect is it's great for us humans often to see these steps.
另一个好的副作用是,对我们人类来说,经常看到这些步骤是很好的。
side effect 常用搭配
副作用;附带效果
指主要结果之外附带产生的影响,可好可坏。
It builds trust. But also we learn we can double check things.
它能建立信任。但我们也学到我们可以复核事情。
double check 常用搭配
复核;再次检查
指为确保准确而再次检查某事。
There's a lot in here. I think some of the debate, there's been a lot of debate this year on if the language models like these aha.
这里面有很多内容。我认为今年有很多关于像这些语言模型是否具有顿悟时刻的争论。
I think the aha moments are kind of fake because in pre-training, you essentially have seen the whole Internet.
我认为这些顿悟时刻有点虚假,因为在预训练中,你基本上已经看过了整个互联网。
So you have definitely seen people explaining their work, even verbally, like a transcript of a math lecture.
所以你肯定见过人们解释他们的工作,甚至口头解释,就像数学讲座的文字记录一样。
You try this. Oh, I messed this up.
你试试这个。哦,我搞砸了。
messed this up 常用搭配
把这件事搞砸了、弄错了
口语中承认自己犯错或操作失败时使用,语气轻松。
And what reinforcement learning is this RLVR is very good at doing is amplifying these behaviors
而强化学习,也就是这个RLVR,非常擅长做的就是放大这些行为
is very good at doing is 句型
非常擅长做的事情就是……
[subject] is very good at doing is [verb phrase]
用于强调某事物最擅长的具体方面,后接动词原形。
because they're very useful in enabling the model to think longer and to check its work.
因为它们非常有用,能让模型思考更长时间并检查自己的工作。
check its work 常用搭配
检查自己的工作、核对结果
指完成任务后回头验证是否正确,常用于学习或工作场景。
And I agree that it is very beautiful
我同意这非常美妙
that this training kind of the model learns to amplify this in a way that is just so useful at the final answers being better.
这种训练让模型学会以某种方式放大这一点,这种方式对于最终答案变得更好非常有用。
I can give you also a hands-on example.
我也可以给你一个实际例子。
hands-on example 常用搭配
实际操作的例子、亲身实践的例子
用于引出具体的、可动手验证的实例,区别于理论说明。
I was training the Gwent 3 base model with RLVR on Math 500.
我当时在用RLVR在Math 500上训练Gwent 3基础模型。
The base model had an accuracy of about 15%.
基础模型的准确率大约为15%。
Just 50 steps, like in a few minutes with RLVR, the model went from 15% to 50% accuracy.
仅仅50步,就像几分钟内使用RLVR,模型准确率从15%提升到了50%。
And you can't tell me it's learning anything fundamentally about math.
你无法告诉我它从根本上学到了任何关于数学的东西。
you can't tell me 地道口语
你没法说服我、我不相信
口语中表达强烈质疑或反驳对方观点时使用。
The Quinn example is weird because there's been two papers this year,
Quinn这个例子很奇怪,因为今年已经有两篇论文了,
one of which I was on that talks about data contamination in Quinn,
其中一篇我参与了,讲的是Quinn中的数据污染,
and specifically that they train on a lot of this special mid-training phase that we just have like a minute on because it's weird.
特别是他们在很多这种特殊的中间训练阶段上训练,而我们只有一分钟左右,因为那很奇怪。
So they train on problems that are almost identical to math.
所以他们在几乎和数学题一模一样的问题上训练。
Exactly. And so you can see that basically the RL, it's not teaching the model any new knowledge about math.
没错。所以你可以看到,基本上这个强化学习并没有教给模型任何关于数学的新知识。
You can't do that in 50 steps.
你不可能在50步里做到这一点。
So the knowledge is already there in the pre-training. You're just unlocking it.
所以知识在预训练中就已经存在了。你只是在解锁它。
I still disagree with the kind of premise because there's a lot of weird complexities that you can't prove.
我还是不同意这个前提,因为有很多奇怪的复杂性你无法证明。
Because one of the things that points to weirdness is that if you take the Quen3 so-called base model
因为指向奇怪之处的一点是,如果你拿Quen3所谓的基础模型
and you can you could Google on the screen, you could Google like math data set hugging face and you could take a problem.
你可以在屏幕上谷歌,你可以谷歌像math数据集hugging face,然后你可以拿一道题。
And what you do if you put it into Quen3 base, all these math problems have words.
如果你把它放进Quen3基础模型,所有这些数学题都有文字。
So it'd be like Alice has five apples and takes one and gives three to whoever.
所以会像Alice有五个苹果,拿走一个,然后给某人三个。
There are these word problems with these Qwen-based models.
这些基于Qwen的模型存在这些文字问题。
Why people are suspicious of them is if you change the numbers but keep the words, Qwen will produce like a very high without tools.
人们怀疑它们的原因是,如果你改变数字但保留文字,Qwen在没有工具的情况下会产生非常高的结果。
Will produce a very high accuracy, like decimal representation of the answer, which means there's some like at some time it was shown problems that were almost identical to the test set.
会产生非常高的准确率,比如答案的十进制表示,这意味着有时它被展示的问题几乎与测试集相同。
And it was using tools to get a very high precision answer, but a language model without tools will never actually have this.
它使用工具来获得非常高精度的答案,但没有工具的语言模型永远不会真正拥有这一点。
So it's kind of been this big debate in the research community is like how much of these reinforcement learning papers that are training on Qwen and measuring specifically on this like math benchmark where there's been multiple papers talking about contamination is like how much can you believe them?
所以这有点像研究界的一个大争论,就是这些在Qwen上训练并专门在这个数学基准上测量的强化学习论文,已经有多篇论文讨论污染问题,就像你能多大程度上相信它们?
And I think this is what caused the reputation of RLVR being about formatting because you can get these gains so quickly and therefore it must already be in the model.
我认为这就是导致RLVR被认为只是关于格式化的原因,因为你可以如此快速地获得这些收益,因此它一定已经在模型中了。
But there's a lot of complexity here that it's not really like controlled experimentation,
但这里有很多复杂性,它并不真的像是受控实验,
so you don't really know. But if it weren't true, I would say distillation wouldn't work, right?
所以你并不真的知道。但如果这不是真的,我会说蒸馏就不会起作用,对吧?
I mean, distillation can work to some extent, but the thing is that is, I think, the biggest problem in LM research: this contamination,
我的意思是,蒸馏在一定程度上可以起作用,但问题是,我认为这是语言模型研究中最大的问题:这种污染,
because we don't know what's in the data. It's unless you have a new data set, it's really impossible.
因为我们不知道数据里有什么。除非你有一个新的数据集,否则真的不可能。
And the same, you mentioned, um, math, the math data set, which is, if a question and an answer and an explanation is given,
同样地,你提到,嗯,数学,那个数学数据集,就是如果给出一个问题、一个答案和一个解释,
but then also even something simpler like MMLU, which is a multiple choice benchmark,
但还有更简单的,比如MMLU,这是一个多项选择基准测试,
if you just change the format slightly, um, like I don't know, you use a dot instead of a parenthesis or something like that,
如果你只是稍微改变一下格式,嗯,比如我不知道,你用点号代替括号之类的,
the model accuracy will vastly differ. I think that that could be like a model issue rather than a general issue.
模型的准确率会大不相同。我认为那可能是一个模型问题,而不是一个普遍问题。
It's not even malicious by the developers of the LM, like, hey, we want to cheat at that benchmark.
这甚至不是语言模型开发者的恶意行为,比如,嘿,我们想在那个基准测试上作弊。
It's just, it has seen something at some point. And I think the only fair way to evaluate an LLM is to have a new benchmark that is after the cutoff date,
只是它在某个时候见过一些东西。我认为评估LLM唯一公平的方式是有一个在截止日期之后的新基准测试,
when the LLM was deployed.
当LLM部署时。
Can we lay out what would be the sort of the recipe of all the things that will be going to post-training?
我们能否列出进入后训练阶段的所有事项的配方?
lay out 常用搭配
列出、阐述清楚
用于把计划、步骤或想法有条理地说明出来。
And you mentioned our RLVR was a really exciting, effective thing.
你提到我们的RLVR是一件非常令人兴奋且有效的事情。
Maybe we should elaborate.
也许我们应该详细说明。
RLHF still has a really important component to play.
RLHF仍然有一个非常重要的组成部分要发挥。
has a really important component to play 句型
仍然有很重要的部分要发挥(作用)
[subject] has a really important component to play
用于强调某事物在整体中仍扮演重要角色,后常接 in 说明领域。
What kind of other ideas are there on post-training?
在后训练方面还有哪些其他想法?
I think you can kind of take this in order.
我认为你可以按顺序来看这个。
I think you could view it as what made 01, which is this first reasoning model possible,
我认为你可以把它看作是什么使01成为可能,这是第一个推理模型,
or what will the latest model be.
或者最新的模型会是什么。
And they actually have, you're going to have similar interventions at these where
他们实际上有,你会在这些地方有类似的干预,其中
you start with mid-training and the thing that is rumored to enable 01 and similar models is really careful data curation
你从中期训练开始,而据传使01和类似模型成为可能的是非常仔细的数据整理
where you're providing a broad set of like what is called reasoning traces,
你提供广泛的所谓推理轨迹,
which is just the model generating words in a forward process that is reflecting like breaking down a problem into intermediate steps and trying to solve them so
这只是模型在前向过程中生成词语,反映为将问题分解为中间步骤并尝试解决它们,所以
at mid-training you need to have data that is similar to this to make it so that when you move into post-training primarily with this verifiable rewards it can learn and then what is happening today is you're figuring
在中期训练时,你需要有类似这样的数据,以便当你主要使用这种可验证奖励进入后训练时,它可以学习,而今天正在发生的事情是你正在弄清楚
out which problems to give the model and how long you can train it for
找出该给模型哪些问题,以及你能训练它多久
and like how much inference you can enable the model to use when solving these verifiable problems
以及你能让模型在解决这些可验证问题时使用多少推理
so as models get better, certain problems are no longer like the model will solve them 100 of the time
所以随着模型变得更好,某些问题就不再是模型能百分之百解决的了
and therefore there's very little signal in this if we pull, if we look at the GRPO equation, this one is famous for this
因此这里面几乎没有信号,如果我们看GRPO方程,这个方程就是因此而出名的
because essentially the reward given to the agent is based on how good a given action is relative to the other answers to that same problem
因为本质上给智能体的奖励取决于某个给定动作相对于同一问题的其他答案有多好
so if all the problems get the same answer, there's no signal in these types of algorithms
所以如果所有问题都得到相同的答案,这类算法里就没有信号
so what they're doing is they're finding harder problems, which is why you hear about things like scientific domains
所以他们做的就是寻找更难的问题,这就是为什么你会听到科学领域之类的东西
which is like, that's so hard, like getting anything right there, if you have a lab or something, it just generates so many tokens, or much harder software problems
这就像是,那太难了,在那里做对任何事都很难,如果你有一个实验室之类的,它就会生成非常多的token,或者难得多的软件问题
so the frontier models are all pushing into these harder domains and they can train on more problems and the model will learn more skills at once
所以前沿模型都在向这些更难的领域推进,它们可以在更多问题上训练,模型就能一次学到更多技能
the RLHF link to this is kind of like RLHF has
与此相关的RLHF联系有点像,RLHF一直
been and still is kind of like the finishing touch on the models
一直是,而且现在仍然是模型上的点睛之笔
the finishing touch 常用搭配
最后的点睛之笔、收尾完善
指在最后阶段添加的使整体更完美的部分。
where it makes the models more useful by improving the organization or style or
它通过改进组织或风格或
or tone there's different things that resonates to different audiences
或语气,让模型更有用,不同的事物能引起不同受众的共鸣
like some people like a really quirky model and RLHF could be good at enabling that personality
比如有些人喜欢非常古怪的模型,而RLHF可能擅长实现那种个性
and some people hate this like markdown bulleted list thing that the models do
而有些人讨厌模型做的这种Markdown项目符号列表
but it's actually really good for quickly parsing information
但它实际上非常有利于快速解析信息
and RLHF this human feedback stage is really great for just give putting this into the model at the end of the day
而RLHF这个人类反馈阶段非常擅长在一天结束时把这些放入模型中
so it's what it made ChatGPT so magical for people
所以这就是它让ChatGPT对人们如此神奇的原因
and that use has actually remained fairly stable this formatting
而这种使用实际上一直相当稳定,这种格式
can also help the models get better at math problems, for example.
也可以帮助模型更好地解决数学问题,例如。
So it's like the border between style and formatting
所以这就像风格和格式之间的边界
and like the method that you use to answer a problem is actually, they're all very closely linked in terms of when you're training these models,
而就像你用来回答问题的方法实际上,它们在你训练这些模型时都非常紧密地联系在一起,
which is why ROHF can still say make a model better at math,
这就是为什么ROHF仍然可以说让模型更擅长数学,
but these verifiable domains are a much more direct process to doing this
但这些可验证领域是一个更直接的过程来做这件事
because it's kind of makes more sense with the problem formulation,
因为它有点更符合问题表述,
which is why it kind of ends up all forming together.
这就是为什么它最终都融合在一起。
But to summarize, it's like mid-training is give the model the skills it needs to then learn.
但总结一下,这就像中期训练是给模型它之后学习所需的技能。
RL and verifiable rewards is let the model try a lot of times.
RL和可验证奖励是让模型尝试很多次。
So put a lot of compute into trial and error learning across hard problems.
所以把大量计算投入到解决难题的试错学习中。
trial and error 常用搭配
试错
通过反复尝试并从错误中学习来解决问题时使用。
And then RLHF would be like, finish the model, make it easy to use, and kind of just round the model out.
然后RLHF就像是,完成模型,让它易于使用,并稍微完善一下模型。
round the model out 常用搭配
使模型更完善、更全面
表示对某事物做最后的补充完善,使其更完整好用。
Can you comment on the amount of compute required for RLVR?
你能评论一下RLVR所需的计算量吗?
It's only gone up and up.
它只会不断上升。
gone up and up 常用搭配
不断上升、持续增长
口语中强调某数量持续不断地增加。
So I think Grok4 was famous for saying they use a similar amount of compute for pre-training and post-training.
所以我认为Grok4以声称他们在预训练和后训练中使用相似的计算量而闻名。
Back to the scaling discussion, they involve very different hardware for scaling.
回到扩展讨论,它们涉及非常不同的硬件来进行扩展。
Pre-training is very compute bound, which is like this flops discussion, which is just how many matrix multiplications can you get through in one time.
预训练非常受计算限制,就像这个FLOPs讨论,即一次能完成多少次矩阵乘法。
And because RL, you're generating these answers, you're trying the model in the real world environments.
而因为RL,你在生成这些答案,你在真实世界环境中尝试模型。
It ends up being much more memory bound because you're generating long sequences and the attention mechanisms have this behavior where you get a quadratic increase in memory as you're getting to longer sequences.
它最终更受内存限制,因为你在生成长序列,而注意力机制有这种行为:随着序列变长,内存会呈二次增长。
So the compute becomes very different.
所以计算变得非常不同。
So when in pre-training, we would talk about a model.
所以在预训练中,我们会谈论一个模型。
I think if we go back to the Biden administration executive order, it's like 10 to the 25th flops to train a model.
我认为如果我们回到拜登政府的行政命令,训练一个模型需要10的25次方FLOPs。
If you're using flops in post-training, it's a lot weirder because the reality is just like, how many hours are you allocating how many GPUs for?
如果你在后训练中使用浮点运算,那就奇怪多了,因为现实情况就像,你分配了多少小时、多少GPU?
And I think in terms of time, the RL compute is getting much closer because you just can't put it all into one system.
我认为在时间方面,强化学习的计算量正变得越来越接近,因为你不能把所有东西都放进一个系统里。
Like pre-training is so computationally dense where all the GPUs are talking to each other and it's extremely efficient.
就像预训练的计算密度非常高,所有GPU都在互相通信,而且效率极高。
where RL has all these moving parts and it can just take a long time to generate a sequence of 100,000 tokens.
而强化学习有所有这些动态部分,生成一个10万token的序列可能需要很长时间。
moving parts 常用搭配
(系统中)不断变化、需要协调的组成部分
形容一个系统复杂、有多个动态环节需要处理。
If you think about GPT 5.2 Pro taking an hour, it's like, what if your training run has a sample for an hour and you have to make it so that's handled efficiently?
如果你想想GPT 5.2 Pro花一个小时,就像,如果你的训练运行有一个样本需要一个小时,你必须让它被高效处理?
So I think in GPU hours or just like wall clock hours, the RL runs are probably approaching the number of days as pre-training,
所以我认为在GPU小时或实际时钟时间上,强化学习的运行时间可能正接近预训练的天数,
but they probably aren't using as many GPUs at the same time.
但它们可能没有同时使用那么多GPU。
There's rules of thumb where in labs, it's like you don't want your pre-training runs to last more than like a month because they fail catastrophically.
有一些经验法则,在实验室里,就像你不希望你的预训练运行持续超过一个月,因为它们会灾难性地失败。
rules of thumb 常用搭配
经验法则
指基于经验而非精确计算的实用准则。
And if you were planning a huge cluster to be held for two months and then it fails on day 50, the opportunity costs are just so big.
如果你计划一个大型集群运行两个月,然后它在第50天失败了,机会成本就太大了。
So you kind of don't want to just,
所以你不想要只是,
people don't want to put all their eggs in one basket,
人们不想把所有鸡蛋放在一个篮子里,
put all their eggs in one basket 地道口语
把所有希望押在一件事上,孤注一掷
劝人不要只依赖单一方案或资源,以免失败时全盘皆输。
which is like GPT-4 was like the ultimate YOLO run
这就像GPT-4是终极的YOLO运行
YOLO run 地道口语
孤注一掷的尝试(YOLO 即 you only live once)
口语中形容冒险、不计后果的一次性大胆行动。
and nobody ever wanted to do it before
以前没有人想这样做
where it took like three months to train
训练需要三个月
and everybody was shocked that it worked
每个人都震惊它成功了
where I think people are a little bit more cautious and incremental now.
我认为现在人们更加谨慎和渐进。
So RL, VR is more, let's say, unlimited how much you can train or get still benefit
所以RL、VR更像是,比如说,你可以训练或获得收益的量是无限的
where rlhf because it's a preference tuning it you reach a certain point where it doesn't really make sense to spend more rl budget on that
而RLHF因为它是偏好调优,你达到某个点后,再花更多RL预算就没有意义了
so just a step back with preference tuning
所以退一步说,偏好调优
step back 常用搭配
退一步(从更宏观的角度看)
在讨论中暂时抽离细节,从整体角度重新审视问题时使用。
so there are multiple people that can give multiple let's say explanations for the same thing
所以有多个人可以对同一件事给出多个,比如说,解释
and they can both be correct but at some point you learned a certain style and it doesn't make sense to you know iterate on it
它们可能都是正确的,但在某个时候你学会了一种特定风格,再迭代就没有意义了
my favorite example is like if relatives ask me what laptop they should buy i give them an explanation
我最喜欢的例子是,如果亲戚问我该买什么笔记本电脑,我会给他们一个解释
or ask like yeah what is your um use case like
或者问,比如,是的,你的使用场景是什么
they for example prioritize battery life and storage other
他们例如优先考虑电池寿命和存储,其他
people like us for example we would prioritize ram and compute and so but both both answers are
像我们这样的人,例如,我们会优先考虑内存和计算能力,所以但两个答案都是
Correct, but different people require different answers,
没错,但不同的人需要不同的答案,
and with preference tuning, well, you're trying to average somehow,
而通过偏好调优,嗯,你试图以某种方式取平均,
like you are asking the data labelers to give you the right or not the right, the preferred answer,
就像你让数据标注员给你正确或不是正确、首选的答案,
and then you train on that, but at some point, yeah, you learn that average preferred answer,
然后你基于此训练,但到了某个时候,是的,你学到了那个平均的首选答案,
and there's no, I think, reason to keep training longer on it,
而且我认为没有理由继续训练更久,
because, you know, it's just a style where with our LVR, you literally give the model
因为,你知道,这只是一种风格,用我们的LVR,你直接给模型
well, you let the model solve more and more complex, difficult problems,
嗯,你让模型解决越来越复杂、困难的问题,
and so I think that it makes more sense to allocate more budget long term to our RVR,
所以我认为长期来看,把更多预算分配给我们的RVR更合理,
and also that right now we are in LRVR 1.0 land,
而且现在我们还处在LRVR 1.0阶段,
where it's still like that simple thing where we have a question and answer,
它仍然像那种简单的东西,我们有一个问题和答案,
but we don't do anything with the one stuff in between.
但我们没有对中间的那部分做任何处理。
So there was a, I mean, multiple research papers also by Google, for example,
所以有,我是说,多篇研究论文,比如谷歌的,
on process reward models that also give scores for the explanation,
关于过程奖励模型,它们也为解释打分,
how correct is the explanation, and I think that will be the next thing.
解释有多正确,我认为这将是下一步。
Let's say our LVR 2.0 for this year focusing in between question and answer,
假设我们今年的LVR 2.0专注于问题和答案之间,
like how to leverage that information
比如如何利用这些信息
the explanation to improve the explanation and help it to get better accuracy
改进解释并帮助它获得更好准确性的解释
but then so that that's one angle
但那么所以那是一个角度
And there was a DeepSeq Math version two paper where they also had interesting inference scaling there.
还有一篇 DeepSeq Math 第二版论文,他们在那里也有有趣的推理扩展。
First, they had developed models that grade themselves a separate model.
首先,他们开发了给自己打分的模型,一个单独的模型。
And I think that that will be one aspect.
我认为那将是一个方面。
And the other, like Nathan mentioned, it will be for LRVR branching into other domains.
另一个,正如 Nathan 提到的,将是 LRVR 分支到其他领域。
The place where people are excited are value functions, which is pretty similar.
人们兴奋的地方是价值函数,这非常相似。
So process reward models are kind of like process reward models assign value.
所以过程奖励模型有点像过程奖励模型分配价值。
how good something is to each kind of intermediate step in a reasoning process where value functions
在推理过程中,价值函数对每个中间步骤评估某件事有多好
apply value to every token the language model generates both of these
对语言模型生成的每个 token 应用价值,这两者
have been largely unproven in the language modeling and this reasoning model era people are more optimistic about value functions forever for whatever reason now
在语言建模和这个推理模型时代,这些在很大程度上尚未被证实,现在人们出于某种原因对价值函数永远更加乐观
i think process reward models were tried a lot more in this pre-01 pre-reasoning model era and a lot of people had a lot of headaches with them
我认为过程奖励模型在这个 pre-01 前推理模型时代被尝试得更多,很多人对它们有很多头疼的问题
so i think a lot of it is the human nature of like value models have a very deep history in reinforcement learning
所以我认为很多是人性使然,就像价值模型在强化学习中有着非常深厚的历史
they're one of the first things that were core to like
它们是最早的核心事物之一,就像
Deep reinforcement learning existing is like training value models in this.
深度强化学习现有的就像是在这训练价值模型。
So right now the literature people are excited about trying value models,
所以现在文献中人们很兴奋地尝试价值模型,
but there's very little proof in it, and there are negative examples in trying to scale up process reward models.
但其中证据很少,而且在尝试扩大过程奖励模型方面有负面例子。
These things don't always hold in the future.
这些事情在未来并不总是成立。
I think we came to this discussion by talking about scaling and a simple way to summarize what you're saying
我认为我们通过讨论扩展和一种简单的方式来总结你所说的,才来到这个讨论。
with like you don't want to do too much RLHF,
就像你不想做太多RLHF,
which is eventually the signal scales is people have worked on RLHF for language models for years, especially in intense interest after ChatGPT,
这最终是信号扩展,人们已经在语言模型的RLHF上工作多年,尤其是在ChatGPT之后兴趣浓厚,
and this the first release of a reasoning model trained with RLVR opening eyes 01
而这第一个用RLVR训练的推理模型发布,睁开了眼睛01
had a scaling plot where if you increase the training compute logarithmically, you get a linear increase in evaluations,
有一个扩展图,如果你对数增加训练计算量,评估会线性增加,
and this has been reproduced multiple times. I think DeepSeek had a plot like this,
而且这已经被多次重现。我认为DeepSeek有一个像这样的图,
but there's no scaling lot for RLHF where if you log increase the compute, you get some performance. In fact,
但RLHF没有扩展图,如果你对数增加计算量,你会得到一些性能。事实上,
the seminal scaling paper for RLHF is scaling laws for reward model over optimization.
RLHF的开创性扩展论文是奖励模型过度优化的扩展定律。
So it's like that's a big line to draw with RLVR and the methods we have now and
所以就像用RLVR和我们现在的方法画一条大线,而且
a big line to draw 常用搭配
一条很难划定的界限
用于说明两个概念或做法之间很难明确区分时。
in the future like they will follow the scaling paradigm
在未来,他们会遵循扩展范式
follow the scaling paradigm 常用搭配
遵循扩展范式
用于描述某个领域或团队按照既有的规模化发展路线前进。
which is like the best runs you can let to run for an extra 10x and you get a few x performance
这就像你可以让最好的运行多跑10倍,然后获得几倍的性能提升
a few x performance 常用搭配
几倍的性能提升
口语中表示性能提升的倍数,x 读作 times。
but you can't do this with rlhf and that is just going to be field defining and how people approach them
但你不能用RLHF做到这一点,而这将成为领域定义性的,以及人们如何对待它们
field defining 常用搭配
具有领域定义意义的
形容某事物会决定整个领域未来的走向。
where i'm a shill for people academically to do rlhf and that's a good way to describe it
在这里我是学术人士做RLHF的推销员,这是一个很好的描述方式
a shill for 常用搭配
为……摇旗呐喊的人
口语中略带贬义或自嘲,指替某人或某事鼓吹的人。
it's like to do the best rlhf you might not need the extra 10 or 100x of compute
这就像要做最好的RLHF,你可能不需要额外的10倍或100倍计算资源
but to do the best rlvr you do so i think there's a
但要做最好的RLVR,你就需要,所以我认为有一个
what i say is a seminal paper from what was a meta internship
我说的是来自Meta实习期的一篇开创性论文
seminal paper 常用搭配
开创性论文
学术语境中形容对该领域有奠基性影响的论文。
it's called it's like the art of scaling reinforcement learning with language models
它叫做,就像《用语言模型扩展强化学习的艺术》
they're what they describe as a framework is scale rl and their incremental experiment was like 10 000 b 200 hours
他们描述的框架是扩展强化学习,他们的增量实验大约是10000亿参数200小时
which is like thousands or tens of thousands of dollars per experiment.
这就像每次实验需要数千或数万美元。
And they do a lot of them, which is just like this cost is not accessible to the average academic,
他们做了很多这样的实验,这就像这个成本对普通学术人士来说是无法承受的,
not accessible to 常用搭配
对……来说无法承受/无法获得
用于说明某资源或机会超出某类人的能力范围。
which is a hard equilibrium where it's trying to figure out how to learn from each community.
这是一个艰难的平衡,它试图弄清楚如何从每个社区学习。
I was wondering if we could take at this point a bit of a tangent and talk about education and learning.
我想知道,我们现在能不能稍微跑个题,聊聊教育和学习。
If you're somebody listening to this, who's a smart person interested in programming, interested in AI.
如果你是正在听这个的人,一个对编程感兴趣、对AI感兴趣的聪明人。
So I presume building something from scratch is a good beginning.
所以我想,从零开始构建一些东西是个好的开始。
from scratch 常用搭配
从零开始
指不借助现成基础,从头构建或学习某事物。
So can you just take me through what you would recommend people do?
所以你能带我过一遍你会推荐人们做什么吗?
take me through 常用搭配
带我过一遍
请对方逐步讲解或演示某个过程。
So I would personally start, like you said, implementing a simple model from scratch that you can run on your computer.
所以我个人会像你说的那样,从零开始实现一个简单的模型,可以在你自己的电脑上运行。
The goal is not if you build a model from scratch to have like something you use every day for your personal projects.
目标不是如果你从零开始构建一个模型,就能拥有像你每天用于个人项目的东西。
Like it's not going to be your personal assistant replacing an existing open-weight model or Chachupity.
就像它不会成为你的个人助理,取代现有的开放权重模型或Chachupity。
It's to see what exactly goes into the LLM, what exactly comes out of the LLM, how the pre-training works in that sense, on your own computer preferably.
而是为了看清LLM内部到底输入了什么,LLM到底输出了什么,在这个意义上预训练是如何工作的,最好是在你自己的电脑上。
And then you learn about the pre-training, the supervised fine-tuning, the attention mechanism.
然后你学习预训练、监督微调、注意力机制。
You get a solid understanding of how things work.
你会对事物如何运作有一个扎实的理解。
a solid understanding of 常用搭配
对……有扎实的理解
形容对某主题掌握得牢固可靠。
But at some point you will reach a limit because small models can only do so much.
但在某个时候你会达到一个极限,因为小模型只能做这么多。
can only do so much 常用搭配
能力有限,只能做到这个程度
用于说明某人或某物有其固有局限。
And the problem with learning about LLMs at scale is I would say it's exponentially more complex to make a larger model
而大规模学习LLM的问题在于,我想说,做一个更大的模型要复杂得多,是指数级的复杂。
because it's not that the model just becomes larger, you have to now think about sharding your parameters across multiple GPUs.
因为并不是模型只是变大了,你现在还得考虑把参数分片到多个GPU上。
Even for the KV cache, there are multiple ways you can implement it.
即使是KV缓存,也有多种实现方式。
One is just to understand how it works, just to grow the cache.
一种是只理解它的工作原理,只是去增长缓存。
It's like a cache you grow step by step by, let's say, concatating lists, growing it.
这就像一个缓存,你一步一步地增长,比如说,拼接列表,让它增长。
But then that wouldn't be optimal in GPUs. You wouldn't do that.
但那样在GPU上不是最优的。你不会那么做。
You would pre-allocate a tensor and then fill it in.
你会预先分配一个张量,然后再填充它。
But that adds, again, another 20, 30 lines of code.
但那又增加了20、30行代码。
And for each thing, you add so much code.
而且每件事你都要加这么多代码。
And I think the trick with the book is basically to understand how the LLM works.
我觉得这本书的诀窍基本上就是理解LLM是如何工作的。
the trick with 常用搭配
……的诀窍
用于引出做某件事的关键方法或要点。
It's not going to be your production level LLM.
它不会成为你的生产级LLM。
But once you have that, you can understand the production level LLM.
但一旦你掌握了那个,你就能理解生产级的LLM。
So you're trying to always build an LLM that's going to fit on one GPU.
所以你总是试图构建一个能装在一块GPU上的LLM。
Yes.
是的。
Most of them I have, I have some bonus materials on some MOE models.
大部分我都有,我有一些关于MOE模型的额外材料。
I think one or two of them, they may require multiple GPUs, but the goal is to have it on one GPU.
我觉得其中一两个可能需要多块GPU,但目标是让它在一块GPU上运行。
And the beautiful thing is also you can self-verify.
而美妙之处在于,你还可以自我验证。
It's almost like RLVR when you code these from scratch.
当你从零开始编写这些代码时,这几乎就像RLVR一样。
You can take an existing model from the Hugging Face Transformer library.
你可以从Hugging Face Transformer库中拿一个现有的模型。
So the Hugging Face Transformer library is great,
所以Hugging Face Transformer库很棒,
but if you want to learn about LLMs, I think that's not the best place to start
但如果你想学习LLM,我觉得那不是最好的起点,
because the code is so complex because it has to fit so many use cases.
因为代码太复杂了,因为它必须适配那么多使用场景。
Also, some people use it in production.
另外,有些人在生产环境中使用它。
It has to be really sophisticated and it's really intertwined and really hard.
它必须非常精密,而且真的相互交织,非常难。
It's not linear to read.
读起来不是线性的。
It was started as a fine-tuning library and then it grew to be like the standard representation
它最初是一个微调库,后来发展成了标准表示方式,
of every model architecture and the way it is loaded.
用于每种模型架构以及它的加载方式。
So Hugging Face is like the default place to get a model
所以Hugging Face就像是获取模型的默认地方,
and Transformers is the software that enables it so people can easily load a model and do something basic with it.
而Transformers是让它成为可能的软件,这样人们就能轻松加载模型并用它做一些基本的事情。
And all Frontier labs that have open weight models have a Hugging Face Transformers version of it,
所有拥有开放权重模型的前沿实验室都有一个Hugging Face Transformers版本,
like from DeepSeq to GPT-OSS.
比如从DeepSeq到GPT-OSS。
That's like the canonical weight that you can load there.
那就像是你可以加载的规范权重。
But again, also even Transformers, the library is not used in production.
但再说一次,即使是Transformers这个库,也不用于生产环境。
People use then SGLang or VLLM and it adds another layer of complexity.
人们用的是SGLang或VLLM,这又增加了一层复杂性。
We should say that the Transformers library has like 400 models.
我们应该说,Transformers库大概有400个模型。
So it's one library that tries to implement a lot of LLMs.
所以它是一个试图实现很多大语言模型的库。
And so you have a huge code base, basically.
所以基本上你有一个巨大的代码库。
It's like huge. It's like, I don't know, maybe millions, hundreds of thousands of lines of code.
它非常大。就像,我不知道,也许几百万、几十万行代码。
And it's like understanding the part that you want to understand is finding the needle in the haystack.
而理解你想理解的那部分,就像大海捞针。
finding the needle in the haystack 地道口语
大海捞针
形容在大量信息中寻找极难找到的目标。
But what's beautiful about it is you have a working implementation.
但它的美妙之处在于你有一个可运行的实现。
And so you can work backwards from it.
所以你可以从它倒推回去。
work backwards from 常用搭配
从……倒推
指从结果或成品出发,反向推导其原理或过程。
what I would recommend doing or what I also do is if I want to understand, for example,
我建议做的,或者我自己也会做的,是如果我想理解,比如说,
how almost three is implemented, I would look at the weights in the model hub, the config file.
几乎三(模型)是如何实现的,我会去看模型中心里的权重,那个配置文件。
And then you can see, oh, they used so many layers.
然后你就能看到,哦,他们用了这么多层。
They use, let's say, group query attention or multi-head attention in that case.
他们用了,比如说,分组查询注意力,或者在这种情况下是多头注意力。
And you see all the components in like a human readable, I don't know, 100 lines of config file.
然后你能看到所有组件,就像在一个人类可读的,我不知道,100行的配置文件里。
And then you start, let's say, with your GPT-2 model and add these things, you know.
然后你开始,比如说,用你的GPT-2模型,然后加上这些东西,你知道。
And the cool thing here is you can then load the pre-trained weights and see if they work in your model.
而这里很酷的一点是,你可以加载预训练的权重,看看它们在你的模型里能不能用。
And you want to match the same output that you get with a transformer model.
而你想要匹配你用Transformer模型得到的相同输出。
And then you can use that as a, basically as a verifiable reward to make your architecture correct.
然后你可以把它当作一个,基本上当作一个可验证的奖励,来让你的架构正确。
And then it's kind of, sometimes it takes me a day to, with Alma 3, the challenge was rope for the position embeddings.
然后它有点,有时候我要花一天时间,用Alma 3,挑战在于位置嵌入的rope。
They had a yarn extension and there was some custom scaling there and I couldn't quite match these things.
他们有一个yarn扩展,那里有一些自定义缩放,我没能完全匹配这些东西。
And in this struggle, you kind of understand things.
而在这场挣扎中,你会有点理解这些东西。
But the cool thing is, at the end, you know you have it correct
但很酷的一点是,到最后,你知道它是正确的,
because you can unit test it. You can check against the reference implementation.
因为你可以对它做单元测试。你可以对照参考实现来检查。
check against 常用搭配
对照……检查
指用某个参照物来验证结果是否正确。
And I think that's maybe one of the best ways to learn, really, like to basically reverse engineer something.
而且我觉得这也许是最好的学习方式之一,真的,基本上就是去逆向工程某个东西。
reverse engineer 常用搭配
逆向工程
指通过分析成品来推导其设计或实现原理。
I think that that is something that everybody that's interested in getting to AI today should do.
我觉得这是今天每个有兴趣进入AI领域的人都应该做的事。
interested in getting to 常用搭配
有兴趣进入(某个领域)
用于谈论对进入某个行业或领域感兴趣,口语中常见。
And I think that's why I liked your book, is like I came to language models from this RL and robotics field.
我觉得这就是我喜欢你这本书的原因,就像我是从强化学习和机器人领域来到语言模型的。
came to 常用搭配
从……领域转过来/进入
表示自己是从另一个领域进入当前领域的,常用于介绍背景。
Like I never had taken the time to just like learn all the fundamentals.
就像我从来没有花时间去好好学习所有的基础知识。
taken the time to 常用搭配
花时间去做某事
表示抽出时间认真做某事,常用于承认自己之前没做。
And this transformer architecture I describe as being like, so fundamental as like deep learning was a thing that I had to learn in the past
而我把这个 Transformer 架构描述为就像,如此基础,就像深度学习是我过去必须学习的东西一样
and people need to do this. I think that where a lot of people kind of get overwhelmed
而人们需要做这件事。我认为很多人会感到不知所措的地方
get overwhelmed 常用搭配
感到不知所措、被压垮
形容面对太多信息或任务时感到应付不过来。
is how do I apply this to have impact or find like a career path
就是,我该如何应用这个来产生影响,或者找到一条职业道路
apply this to 常用搭配
把某事物应用到……上
用于讨论如何将知识或技能用于实际场景。
because like AI and language models make this fundamental stuff so accessible and people with motivation will learn it
因为像 AI 和语言模型让这些基础知识变得如此容易获取,有动力的人就会去学它
and then it's like how do I get the cycles on goal to contribute to research
然后就是,我该如何把精力投入到目标上,为研究做出贡献
contribute to 常用搭配
为……做出贡献
常用于谈论为研究、项目或社区出力。
and I think that I'm actually fairly optimistic in this
而我认为我其实对此相当乐观
because the field moves so fast that a lot of times the best people don't fully solve a problem
因为这个领域发展得如此之快,以至于很多时候最优秀的人并不会完全解决一个问题
because there's a bigger problem to solve that's very low-hanging fruit, so they move on.
因为有一个更大的问题要解决,而那又是非常容易摘取的果实,所以他们就继续前进了。
low-hanging fruit 地道口语
容易实现的目标、唾手可得的成果
比喻容易取得进展或解决的问题,常用于工作或研究语境。
move on 常用搭配
继续前进、转向下一个
表示不再停留于当前事物,转而做别的事。
And I think that a lot of what I was trying to do in this RLHF book is take post-training techniques
而我认为我在这本 RLHF 书里试图做的很多事情,就是拿后训练技术
and just describe how people think about them influencing the model and what people are doing.
然后只是描述人们如何看待它们影响模型,以及人们在做什么。
And then it's remarkable how many things I just think are just like people stop studying them or don't.
然后很值得注意的是,有多少东西我觉得就是人们停止研究它们,或者没有。
So I think people trying to get narrow after doing the fundamentals is good.
所以我认为人们在打好基础之后尝试变得专精是好事。
And then reading the relevant papers and being engaged in the ecosystem.
然后阅读相关论文并参与这个生态系统。
engaged in 常用搭配
参与、投身于
表示积极参与某个活动或社群。
It's like you actually, the proximity that random people have online from the leading researchers,
就像你实际上,网上随机的人与顶尖研究者的接近程度,
like no one knows who the anonymous account on X and ML is very popular for whatever reason.
比如没人知道X和ML上的匿名账号是谁,不管什么原因它非常受欢迎。
And no one knows who all these people are.
而且没人知道所有这些人都谁。
Like it could just be random people that study the stuff deeply, especially with the AI tools
比如可能只是深入研究这些东西的随机的人,尤其是用AI工具
and just be like, I don't understand this.
然后就说,我不理解这个。
Keep digging into it. I think is a very useful thing.
继续深入挖掘。我觉得这是非常有用的。
digging into 常用搭配
深入钻研、挖掘
表示持续深入地研究某个问题或话题。
But there's a lot of research areas that just like are maybe three papers that you need to read.
但有很多研究领域可能只需要读三篇论文。
And then one of the authors will probably email you back.
然后其中一位作者可能会给你回邮件。
But you have to put a lot of effort into these emails to understand the field.
但你必须在这些邮件上投入很多努力来理解这个领域。
put a lot of effort into 常用搭配
在……上投入很多努力
用于强调为某事付出大量精力。
Like I think it would be for a newcomer easily weeks of work to feel like they can truly grasp like what is a very narrow area.
就像我觉得对一个新手来说,可能需要好几周的工作才能感觉自己真正掌握一个非常狭窄的领域。
But I think going narrow after you have the fundamentals be very useful to people because it's like I became very interested in.
但我认为在掌握基础知识后转向狭窄领域对人们非常有用,因为就像我变得非常感兴趣。
Character training, which is like how you make the model funny or sarcastic or serious,
角色训练,也就是你如何让模型变得有趣、讽刺或严肃,
and like what do you do to the data to do this.
以及比如你要对数据做什么来实现这一点。
And it's like a student at Oxford reached out to me.
就像牛津的一个学生联系了我。
reached out to 常用搭配
主动联系
表示主动联系某人,常用于寻求帮助或合作。
It's like, hey, I'm interested in this, and I advised him, and I was like, that paper now exists.
就像,嘿,我对这个感兴趣,我指导了他,我说,那篇论文现在已经存在了。
And it's like, I don't know, there's like two or three people in the world that were very interested in this.
就像,我不知道,世界上大概只有两三个人对这个非常感兴趣。
He's a PhD student, which gives you an advantage,
他是个博士生,这给了你一个优势,
but like for me, that was a topic I was waiting for someone to be like, hey, I have time to spend cycles on this,
但对我来说,那是一个我一直在等有人来说,嘿,我有时间在这上面花精力的主题,
spend cycles on 地道口语
在……上花精力/时间
口语中表示把精力或时间投入到某件事上。
and I'm sure there's a lot more very narrow things, or you're just like, oh, it doesn't make sense that there was no answer to this,
而且我确信还有更多非常狭窄的领域,或者你只是觉得,哦,这说不通,这个问题竟然没有答案,
doesn't make sense 常用搭配
说不通、不合理
用于表达某事逻辑上不合理或难以理解。
and I think that it's just like there's so much information coming that people are like, I can't grab onto any of these.
我觉得就像有太多信息涌来,人们会觉得,我什么都抓不住。
grab onto 常用搭配
抓住、把握住
比喻理解或掌握某个信息或概念。
But if you just actually stick in an area, I think there's a lot of interesting things to learn.
但如果你真的坚持在一个领域里,我觉得有很多有趣的东西可以学。
stick in 常用搭配
坚持在(某个领域)
表示持续专注于某个领域而不轻易改变。
Yeah, I think you can't try to do it all, because it would be very overwhelming, and you would burn out if you try to keep
是的,我觉得你不能试图什么都做,因为那会非常让人不堪重负,如果你试图跟上一切,你会精疲力竭。
burn out 常用搭配
精疲力竭、累垮
形容因过度劳累而失去精力或动力。
up with everything for me, for example,
对我来说跟上所有东西,比如,
I haven't kept up with computer vision a long time, just focused on LMS,
我很长时间没有跟上计算机视觉了,只是专注于LMS,
kept up with 常用搭配
跟上、保持了解
表示持续关注某个领域的最新进展。
but coming back to your book, for example,
但回到你的书,比如,
I think this is also a really great book and a really good bang for the buck,
我觉得这也是一本非常好的书,性价比很高,
bang for the buck 地道口语
性价比高
口语中表示花同样的钱得到更多价值。
because you want to learn about RLHF,
因为你想了解RLHF,
I wouldn't go out there and read RLHF papers,
我不会去外面读RLHF论文,
because I would be, you would be spending two years contradict,
因为我将会,你将会花两年时间矛盾,
there's, I just edited the book and I was like, there's a chapter where
有,我刚编辑了这本书,我想,有一章
I had to be like, X papers say one thing and X papers say another thing,
我不得不像,X论文说一件事,X论文说另一件事,
and we'll see what comes out to be true.
我们会看看什么最终是真的。
comes out to be 常用搭配
最终结果是、被证明是
表示经过一段时间后事情最终呈现出的结果。
What are some of the just to go through some of the table content,
有哪些只是过一遍表格内容,
some of the ideas we might have missed in the bigger picture of the post-training,
我们可能在后训练的大局中错过的一些想法,
so first of all, you do the problem setup, training overview, what are preferences, preferences, data, and the optimization tools,
所以首先,你做问题设置、训练概述、什么是偏好、偏好、数据,以及优化工具,
reward modeling, regularization, instruction tuning, rejection sampling, reinforcement learning, i.e.
奖励建模、正则化、指令调优、拒绝采样、强化学习,即
policy gradients, direct alignment algorithms, then constitutional AI and AI feedback,
策略梯度、直接对齐算法,然后是宪法AI和AI反馈,
reasoning and inference time scaling, tool use and function calling, synthetic data and distillation,
推理和推理时间扩展、工具使用和函数调用、合成数据和蒸馏,
evaluation, and then open question section over optimization style and information,
评估,然后是优化风格和信息的开放问题部分,
And then product, UX, character, and post-training.
然后是产品、用户体验、角色和后期训练。
So what are some ideas worth mentioning that connect both the educational component and the research component?
那么,有哪些值得提及的想法,既能连接教育部分又能连接研究部分?
You mentioned the character training. It's pretty interesting.
你提到了角色训练。这很有趣。
Character training is interesting because there's so little out of it,
角色训练很有趣,因为从中得到的东西很少,
but we talk about how people engage with these models and we feel good using them
但我们谈论人们如何与这些模型互动,我们使用它们时感觉良好,
because they're positive, but that can go too far. It could be too positive.
因为它们是积极的,但这可能走得太远。它可能过于积极。
go too far 常用搭配
做得过分、走得太远
表示某种做法超出了合适的程度。
And it's like, essentially, it's how do you change your data or decision-making to make it exactly what you want?
本质上,就是如何改变你的数据或决策,使其完全符合你的需求?
And OpenAI has this thing called a model spec, which is essentially their internal guideline for what they want to model to do.
OpenAI有一个叫做模型规范的东西,本质上是他们希望模型做什么的内部指南。
And they publish this to developers. So essentially, you can know what is a failure of OpenAI's training,
他们向开发者发布这个。所以本质上,你可以知道什么是OpenAI训练的失败,
which is like they have the intentions and they haven't met it yet, versus what is something that they actually wanted to do and that you don't like.
这就像他们有意图但尚未实现,而不是他们真正想做的事情而你不喜欢。
And that transparency is very nice.
这种透明度非常好。
But all the methods for curating these documents and how easy it is to follow them is not very well known.
但所有整理这些文档的方法以及遵循它们的容易程度并不广为人知。
I think the way the book is designed is that the Reinforced Learning chapter is obviously
我认为这本书的设计方式是,强化学习那一章显然是
what people want because everybody hears about it with RLVR.
人们想要的,因为每个人都听说过RLVR。
And it's the same algorithms and the same math, but it's just like you can use it in very different documents.
而且算法和数学是一样的,但就像你可以在非常不同的文档中使用它。
So I think the core of RLHF is like how messy preferences are, is essentially rehash of a paper I wrote years ago.
所以我认为RLHF的核心就像偏好有多混乱,本质上是我多年前写的一篇论文的翻版。
But this is essentially the chapter that'll tell you why RLHF is never, ever fully solvable
但这本质上就是那一章,它会告诉你为什么RLHF永远、永远无法完全解决,
because like the way that even RL is set up is that it assumes that preferences can be quantified and that multiple preferences can be reduced to single values.
因为就像即使是RL的设置方式也假设偏好可以被量化,并且多个偏好可以被简化为单一值。
And I think it relates in the economics literature to the von Neumann-Morgenstein utility theorem.
我认为它在经济学文献中与冯·诺依曼-摩根斯坦效用定理相关。
And like that is the chapter where all of that philosophical, economic, and like psychological context, it tells you what gets compressed into doing RLHF.
就像那一章,所有那些哲学、经济以及心理学的背景,它告诉你什么被压缩进了做RLHF中。
So it's like you have all of this, and then later in the book, it's like you use this RL math to make the number go up.
所以就像你有了所有这些,然后在书的后面,就像你使用这个RL数学来让数字上升。
And I think that that's why I think it would be very rewarding for people to do research on,
而且我认为,这就是为什么我觉得人们做这方面的研究会非常有收获,
is because quantifying preferences is something that is just like humans have designed the problem in order to make preferences studyable
是因为量化偏好这件事,就像人类设计出这个问题,就是为了让偏好变得可研究,
but there's kind of fundamental debates on like an example is in a language model response
但存在一些根本性的争论,比如一个例子是,在语言模型的回答中,
you have different things you care about whether it's accuracy or in style
你关心不同的方面,无论是准确性还是风格,
and when you're collecting the data they all get compressed into like i like this more than another and it's like like that is happening
而当你收集数据时,它们都被压缩成类似“我喜欢这个多于另一个”,就像那样的情况在发生,
and there's a lot of philosophical there's a lot of research in other areas of the world that go into like how should you actually do this
而且有很多哲学上的,世界上其他领域有很多研究,涉及你应该如何实际做这件事,
i think social choice theory is the subfield of economics around how you should aggregate preferences.
我认为社会选择理论是经济学中关于应该如何聚合偏好的子领域。
And there's like, I went to a workshop that published a white paper.
而且,比如,我参加了一个发布白皮书的工作坊。
I'm like, how can you think about using social choice theory for RLHF?
我就在想,你怎么能考虑把社会选择理论用于RLHF呢?
So I mostly would want people that get excited about the math to come and have things where they can stumble into and learn this kind of broader context.
所以我主要希望那些对数学感到兴奋的人来,并有一些东西让他们可以偶然接触并学习这种更广泛的背景。
stumble into 常用搭配
偶然接触或无意中发现
用于描述并非刻意计划、而是碰巧遇到某个领域或机会。
I think there's a fun thing. I just keep a list of all the tech reports that I like of reasoning models.
我觉得有件有趣的事。我一直保留着一份清单,列出所有我喜欢的推理模型的技术报告。
So in chapter 14, which is kind of like a short summary of RLVR,
所以在第14章,这有点像是对RLVR的简短总结,
there's just like a gigantic table where I just like list every single reasoning model that I like.
那里就有一个巨大的表格,我就像列出我喜欢的每一个推理模型。
So there's just like, I think in education, a lot of it needs to be like, at this point,
所以就像,我认为在教育中,很多内容在这一点上需要变得像,
it's like what I like because the language models are so good at the math where it's like famous paper,
就像我喜欢的那样,因为语言模型在数学方面非常擅长,就像那篇著名的论文,
direct preference optimization, which is like a much simpler way of solving the problem than RL.
直接偏好优化,这就像是一种比强化学习更简单的解决问题的方法。
The derivations in the appendix skip steps of math.
附录中的推导跳过了数学步骤。
And it's like, I tried for this book, like I redid the derivations and I'm like, what the heck is this log trick that they use to change the math?
就像,我为这本书尝试过,我重新做了推导,然后我想,他们用来改变数学的这个对数技巧到底是什么鬼?
what the heck is 地道口语
……到底是什么鬼
口语中表达困惑或略带不满,用于对某事感到不解时。
But doing it with language models, they're like, this is the log trick.
但用语言模型来做,他们就像,这就是对数技巧。
And I'm like, I don't know if I like this, that the math is so commoditized.
然后我想,我不知道我是否喜欢这样,数学变得如此商品化。
I think some of the struggle in reading this appendix and following the math, I think, is good for learning.
我认为阅读这个附录和跟随数学推导中的一些挣扎,我觉得,对学习是有好处的。
Yeah, so actually, returning to this often, just on the topic of education, you both have brought up the word struggle quite a bit.
是的,所以实际上,经常回到这个话题,就教育而言,你们俩都多次提到“挣扎”这个词。
So there is value.
所以这是有价值的。
If you're not struggling as part of this process, you're not fully following the proper process for learning, I suppose.
如果你没有在这个过程中挣扎,我想你就没有完全遵循正确的学习过程。
Some of the providers are starting to work on models for education, which are designed to not give.
一些提供商开始开发教育模型,这些模型的设计初衷是不直接给出答案。
Actually, I haven't used them, but I would guess they're designed to not give all the information at once and make people work to do this.
其实,我没用过它们,但我猜它们的设计是不一次性给出所有信息,而是让人们努力去完成。
So I think you could train models to do this, and it would be a wonderful contribution.
所以我认为你可以训练模型来做这件事,那将是一个很棒的贡献。
Where all of this stuff in the book, you have to reevaluate every decision for it, which is such a great example.
书里所有这些内容,你都必须重新评估每一个决定,这是一个很好的例子。
I think there's a chance you work on an AI too, which I was like, oh, I think this would be so fun.
我觉得你也有可能在做AI,我当时就想,哦,我觉得这太有趣了。
It makes sense. I do something like that. Did that the other day for video games, for example.
这很合理。我也做过类似的事。比如前几天,我就为电子游戏这么做过。
I sometimes for my pastime play video games. Like I like video games with puzzles, you know, like Zelda and Metroid.
我有时会玩电子游戏作为消遣。比如我喜欢有解谜元素的电子游戏,你知道,像《塞尔达》和《银河战士》。
And there's this new game where I got stuck and I really got stuck. It was okay.
有一款新游戏,我卡住了,真的卡住了。不过还好。
got stuck 常用搭配
卡住了,无法继续
用于描述在游戏、任务或问题中无法推进的状态。
I, you know, I don't want to struggle like two days. And so I use an LLM.
我,你知道,我不想挣扎两天。所以我用了一个大语言模型。
But then you say, hey, please don't add any spoilers. Just, you know, I'm here and there.
但然后你说,嘿,请不要加任何剧透。只是,你知道,我在这里和那里。
What do I have to do next?
我接下来要做什么?
And the same thing you can do, I guess, for math, where you say, okay, I'm here at this point, I'm getting stuck.
我想,数学也是一样,你可以这样做:你说,好,我卡在这个地方了,我卡住了。
Don't give me the full solution, but what is something I could try, you know, like where you kind of carefully probe it.
别给我完整的解法,但我可以尝试什么,你知道,就像你仔细地试探一下。
But the problem here is I think it requires discipline.
但这里的问题是,我认为这需要自律。
And a lot of people do math for like, I mean, a lot of people who enjoy math,
而且很多人做数学是因为,我是说,很多喜欢数学的人,
but there are also a lot of people who need to do it for their homework.
但也有很多人是因为作业才需要做数学。
And then it's like the shortcut.
然后这就成了捷径。
And yeah, we can develop an educational LLM, but the other LLM is still there.
是的,我们可以开发一个教育用的大语言模型,但另一个大语言模型仍然存在。
And there's still a temptation to use the other LLMs.
而且仍然有使用其他大语言模型的诱惑。
I think a lot of people, especially in college, they understand the stuff they're passionate about, they're self-aware about it, and they understand it shouldn't be easy.
我认为很多人,尤其是在大学里,他们理解自己热爱的东西,他们对此有自知之明,他们明白这不应该很容易。
I think we just have to develop a good taste.
我认为我们只需要培养好的品味。
We talk about research taste, school taste, about stuff that you should be struggling on and stuff you shouldn't be struggling on,
我们谈论研究品味、学校品味,谈论你应该挣扎的事情和你不应该挣扎的事情,
which is tricky to know because sometimes you don't have good long-term vision about what would be actually useful to you in your career.
这很难判断,因为有时你对自己职业生涯中真正有用的东西没有好的长期愿景。
But you have to develop that taste.
但你必须培养这种品味。
I was talking to maybe my fiance or friends about this,
我可能跟我的未婚妻或朋友们聊过这个,
and it's like, there's this brief 10-year window where all of the homework and all the exams could be digital.
就像有一个短暂的十年窗口期,所有作业和考试都可以是数字化的。
But before that, everybody had to do all the exams in Blue Book because there was no other way.
但在那之前,所有人都得在蓝皮书上做所有考试,因为别无他法。
And now after AI, everybody's going to need to be in Blue Books and oral exams because everybody could cheat so easily.
而现在有了AI之后,每个人都将需要回到蓝皮书和口试,因为每个人都能轻易作弊。
It's like this brief generation that had a different education system that everything could be digital, but you still couldn't cheat.
就像这短暂的一代人,他们经历了一种不同的教育体系,一切都可以数字化,但你仍然不能作弊。
And now it's just going to go back. It's just pretty funny.
而现在它又要退回去了。这真的挺有趣的。
You mentioned character training.
你提到了角色训练。
Just zooming out on a more general topic, for that topic, how much compute was required?
从更宏观的角度来看,就那个话题而言,需要多少算力?
zooming out 常用搭配
从更宏观的角度来看
用于把讨论从细节拉远到更宏观的层面。
And in general, to contribute as a researcher, are there places where not too much compute is required where you can actually contribute as an individual researcher?
总的来说,作为一名研究者做出贡献,是否存在一些不需要太多算力、你作为独立研究者也能真正做出贡献的领域?
For on the character training thing, I think this research is built on fine-tuning about 7 billion parameter models with LoRA,
关于角色训练这件事,我认为这项研究是基于用LoRA微调约70亿参数的模型,
which is like essentially you only fine-tune a small subset of the weights of the model.
这基本上意味着你只微调模型权重的一小部分。
I don't know exactly how many GPU hours that would take, but it's doable, not doable for every academic.
我不知道具体需要多少GPU小时,但这是可行的,并非每个学术研究者都能做到。
So the situation for some academics is like so dire that the only work you can do is doing inference,
所以对一些学术研究者来说,情况如此严峻,以至于你唯一能做的工作就是进行推理,
where you have closed models or open models and you get completions from them and you can look at them and understand the models and that's very well suited to evaluation,
你可以使用闭源或开源模型,从中获得补全结果,观察并理解这些模型,这非常适合评估,
which you become, you want to be the best at creating representative problems that the models fail on or show certain abilities,
你会成为,你想成为最擅长创建代表性问题的专家,这些问题能揭示模型失败之处或展示某些能力,
which I think that you can break through with this.
我认为你可以通过这个取得突破。
break through 常用搭配
取得突破
用于描述在困难领域中获得重大进展或成功。
So I've like, I think that the top end goal for a researcher working on evaluation, if you want to have career momentum is the frontier labs, pick up your evaluation.
所以我觉得,对于从事评估工作的研究者来说,如果你想要职业发展势头,最终目标就是前沿实验室,拿起你的评估工作。
So it's like, you don't need to have every project do this,
所以就像,你不需要每个项目都这样做,
but if you go from a small university with no compute and you figure out something that Claude struggles with,
但如果你来自一个没有计算资源的小型大学,并且你发现了一些Claude难以处理的事情,
And then the next clod model has it in the blog post.
然后下一个clod模型就在博客文章中提到了它。
There's your career rocket ship.
这就是你职业生涯的火箭飞船。
I think that that's hard, but it's like if you want to scope the maximum possible impact with minimum compute, it's something like that, which is just get very narrow.
我认为这很难,但就像如果你想用最少的计算资源来最大化可能的影响,那就是类似这样的,也就是变得非常狭窄。
And it takes learning of where the models are going.
这需要了解模型的发展方向。
So you need to build a tool that tests where not clod 4.5 will fail.
所以你需要构建一个工具来测试clod 4.5不会在哪里失败。
If I'm going to start a research project, I need to think where the models in eight months are going to be struggling.
如果我要开始一个研究项目,我需要思考八个月后模型会在哪里遇到困难。
But what about developing totally novel ideas?
但是开发全新的想法呢?
this is a trade-off i think that if you're doing a phd you could also be like it's too risky
这是一个权衡,我认为如果你在读博士,你也可以说这太冒险了
to work in language models i'm going way longer term which is like what is what is the thing that's going to define language model development in 10 years which i think that i end up being a person that's pretty practical
从事语言模型工作,我走的是更长远的路,也就是什么是未来10年定义语言模型发展的东西,我认为我最终会成为一个相当务实的人
i mean i went to my phd where it's like i got into berkeley worst case i get a master's and then i
我的意思是我去读博士,就像我进了伯克利,最坏的情况我拿个硕士学位,然后我
go work in tech, it's like I'm very practical about it.
去科技行业工作,我对此非常务实。
So I'm like, the life afforded to people to work at these AI companies,
所以我想,在这些AI公司工作所能获得的生活,
the amount of like OpenAI's average compensation is over a million dollars in stock a year
比如OpenAI的平均薪酬是每年超过一百万美元的股票,
for employee, any normal person in the U.S. to get into this AI lab is transformative for your life.
对员工来说,任何普通美国人能进入这个AI实验室,都会改变你的人生。
So I'm pretty practical of like, there's still a lot of upward mobility working in language models.
所以我很务实,比如,在语言模型领域工作仍然有很多向上流动的机会。
upward mobility 常用搭配
向上流动,晋升机会
用于描述职业或社会地位提升的可能性。
If you're focused and the outcomes is like, look at these jobs,
如果你专注,结果就像,看看这些工作,
but from a research perspective, the transformative of impact in these academic awards,
但从研究的角度来看,这些学术奖项的变革性影响,
that's like be the next Jan LeCun is from not working.
那就像成为下一个Jan LeCun,而不是不工作。
I'm not caring about language model development very much.
我不太关心语言模型的开发。
It's a big financial sacrifice in that case.
那样的话,这是一个巨大的经济牺牲。
So I get to work with some awesome students and they're like,
所以我能和一些很棒的学生一起工作,他们会说,
should I go work in an AI lab?
我应该去AI实验室工作吗?
And I'm like, like you're getting a PhD at a top school or you're going to leave to go to a lab?
我说,比如你在顶尖学校读博士,或者你要离开去实验室?
I'm like, I don't know. Like if you go work at a top lab, I don't blame you.
我说,我不知道。比如如果你去顶尖实验室工作,我不怪你。
I don't blame you 地道口语
我不怪你;我理解你的选择
当别人做出你虽不完全赞同但能理解的决定时,用来表示理解、不责怪对方。
Don't go work at some random startup that might go to zero.
别去那些可能归零的随便的初创公司工作。
go to zero 常用搭配
(公司或价值)归零、彻底失败
口语中形容初创公司或投资可能彻底失败、变得一文不值。
But if you're going to open AI, I'm like, it could be worth leaving a PhD for.
但如果你要去OpenAI,我觉得,这可能值得放弃博士学位。
Let's more rigorously think through this.
让我们更严谨地思考一下。
think through 常用搭配
仔细、全面地思考
表示把一个问题从头到尾认真考虑清楚,常与rigorously、carefully等搭配。
Where would you give a recommendation for people to do a research contribution?
你会推荐人们去哪里做研究贡献?
So the options are academia, so get a PhD, spend five years publishing, compute resources are constrained.
所以选择是学术界,就是读博士,花五年时间发表论文,计算资源受限。
There's research labs that are more focused on open weight models, and so working there.
有些研究实验室更专注于开放权重模型,所以去那里工作。
or closed frontier labs, research labs, OpenAI, Anthropic, XAI, so on.
或者闭源前沿实验室,研究实验室,OpenAI、Anthropic、XAI等等。
The two gradients are the more closed, the more money you tend to get, but also you get less credit.
两个梯度是:越封闭,你往往得到越多钱,但你也得到越少认可。
the more closed, the more money you tend to get 句型
越封闭,你往往得到的钱越多
the more [X], the more [Y] you tend to get
用于表达两个变量同向变化:the more [X], the more [Y],即越……就越……。
So in terms of building a portfolio of things that you've done, it's very clear of what you have done as an academic
所以就建立你做过的事情的作品集而言,作为学者你做了什么非常清楚,
and you have done this, versus if you are going to go trade this fairly reasonable progression for being a cog in the machine, which could also be very fun.
而且你做了这个,而不是如果你要用这个相当合理的进展来换取成为机器中的一颗螺丝钉,这也可能非常有趣。
So I think it's a very different career paths,
所以我认为这是非常不同的职业道路,
but the opportunity cost for being a researcher is very high because PhD students are paid essentially nothing.
但成为研究者的机会成本非常高,因为博士生基本上没有报酬。
opportunity cost 常用搭配
机会成本
经济与决策用语,指为选择某一方案而放弃的其他最佳选择的价值。
So I think it ends up rewarding people that have a fairly stable safety net
所以我觉得,最终受益的是那些有相当稳定安全网的人
safety net 常用搭配
安全网;经济或生活上的保障
指在遇到困难时能提供支持的经济后盾或保障体系。
and they realize that they can operate in the long term,
而且他们意识到自己可以长期运作,
which is they want to do very interesting work and get a very interesting job
也就是他们想做非常有趣的工作,找到一份非常有趣的工作
so it is a fairly like it's a privileged position
所以这相当于是,这是一个有特权的位置
to be like I'm going to see out my PhD and figure it out after
能够说我要读完我的博士,之后再想办法
figure it out 常用搭配
想办法弄明白或解决
表示暂时没有答案或计划,但之后会通过思考、尝试找到办法。
because I want to do this and I think a lot of academic like at the same time
因为我想做这个,而且我觉得很多学术的,同时
the academic ecosystem is getting bombarded by funding getting cut and stuff
学术生态系统正受到经费被削减之类的冲击
so there's just like so many different trade-offs
所以就是有太多不同的权衡取舍
trade-offs 常用搭配
权衡取舍
指在不同选择之间必须放弃一些东西以换取另一些东西的情况。
where I understand plenty of people that are like I can't deal with this funding search
我理解有很多人会说,我受不了这种经费申请
I grant got cut for no reason by the government
我的经费无缘无故被政府砍了
or I don't know what's going to happen
或者我不知道会发生什么
so I think there's a lot of uncertainty and trade-offs that in my opinion favor
所以我觉得有很多不确定性和权衡取舍,在我看来更倾向于
just like take the take the well-paying job with meaningful impact
就是选择那份薪水高、又有意义影响的工作
so it's like not also like you're getting paid to sit around at open AI
所以这也不只是说你拿着工资坐在 OpenAI 里
you're building like the cutting edge of things that are changing
你在打造那些正在改变
the cutting edge 常用搭配
最前沿;领先地位
形容某领域中最先进、最新的技术或发展。
millions of people's relationship to tech
数百万人与科技关系的最前沿
but publication wise they're being more secretive increasingly
但在发表方面,他们越来越保密
So you're publishing less and less and less and less, so you're having a positive impact at scale,
所以你们发表得越来越少,越来越少,所以你们正在产生大规模的积极影响,
at scale 常用搭配
大规模地
表示以很大的规模或范围进行某事,常用于商业和技术语境。
but it's your cog in the machine?
但你是机器里的一个齿轮?
I think it's, honestly, it hasn't changed that much.
我觉得,说实话,这没怎么变。
So I have been in academia. I'm not in academia anymore.
所以我曾经在学术界。我现在不在学术界了。
At the same time, I wouldn't want to miss my time in academia.
同时,我不想错过我在学术界的时光。
But what I wanted to say before I get to that part, I think it hasn't changed that much.
但在我讲到那部分之前,我想说的是,我觉得这没怎么变。
I was working in, like, I was using AI or machine learning methods for applications in computational biology with collaborators.
我当时在做,比如,我和合作者一起使用人工智能或机器学习方法应用于计算生物学。
And a lot of people went from from academia directly to Google.
很多人从学术界直接去了谷歌。
And I think it's the same thing.
我觉得这是一回事。
Back then, the professors were like, you know, sad that their students went into industry because they couldn't carry on their legacy in that sense.
那时候,教授们会,你知道,难过他们的学生进入工业界,因为他们无法以那种方式延续他们的传承。
And I think it's the same thing.
我觉得这是一回事。
It's like it hasn't changed, I think, that much.
就像,我觉得,这没怎么变。
The only thing that has changed is the scale.
唯一改变的是规模。
But, you know, cool stuff was always developed in industry that was closed.
但是,你知道,酷东西一直都是在封闭的工业界开发的。
You couldn't talk about it.
你不能谈论它。
And I think the difference now is, well, your preference.
我觉得现在的区别是,嗯,你的偏好。
Do you like to talk about your work?
你喜欢谈论你的工作吗?
publish or you know you you are more in a closed lab, uh, that's one difference.
发表,或者你知道,你更多是在一个封闭的实验室里,呃,这是一个区别。
The compensation, of course, but it's always been like that, I think. So it really depends on, you know, where you feel comfortable. And it's also nothing is forever. The only thing right now is there's a third option,
当然,薪酬,但我觉得一直都是这样。所以这真的取决于,你知道,你在哪里感到舒服。而且也没有什么是永恒的。现在唯一的事情是还有第三个选择,
nothing is forever 地道口语
没有什么是永恒的
用来安慰或提醒别人,情况总会变化,不必过于执着于当前处境。
which is I'm starting a startup. That's a lot of people doing startups, very risky move, uh, but can be is a high risk, high reward type of situation,
那就是我在创办一家初创公司。很多人都在做初创公司,这是非常冒险的举动,呃,但可能是一个高风险、高回报的情况,
high risk, high reward 常用搭配
高风险,高回报
形容需要承担较大风险但可能获得很大收益的情况。
where joining an industry lab, I think, is pretty safe, you know, also upward mobility. Honestly, I think if once you have been at a industry lab, it will be easier to find future jobs.
而加入一个工业实验室,我认为,是相当安全的,你知道,也有向上流动性。老实说,我认为一旦你曾在工业实验室工作过,将来会更容易找到工作。
upward mobility 常用搭配
向上流动性;晋升机会
指在职业或社会阶层中向上发展的可能性。
But then again, you know, you know, it's like, yeah, how much do you enjoy the team and working on propriety things versus how do you like the publishing work?
但话说回来,你知道,你知道,就像,是的,你有多喜欢这个团队和做专有的事情,相对于你有多喜欢发表工作?
I mean, publishing is stressful. It is, um, you know, like acceptance rate at conferences can be arbitrary, can be very frustrating, but also high reward if you have a paper published, you feel good because your name is on there, you have a high accomplishment, and, you know, I feel like my friends
我的意思是,发表是有压力的。它确实,嗯,你知道,像会议的接受率可能是随意的,可能非常令人沮丧,但也是高回报的,如果你有一篇论文发表,你会感觉很好,因为你的名字在上面,你有很高的成就,而且,你知道,我觉得我的朋友们
who are professors seem on average happier than my friends who work at a frontier lab
教授们平均来说似乎比我在前沿实验室工作的朋友们更快乐
to be totally honest because that's just grounding and the frontier labs definitely do this
说实话,因为那只是基础,而前沿实验室绝对会这样做
996, which essentially is shorthand for work all the time.
996,本质上就是一直工作的简称。
Can you describe 996 as a culture that's, I believe you could say, invented in China and adopted in Silicon Valley?
你能把996描述为一种文化,我相信你可以说,它起源于中国并被硅谷采用吗?
What's 996? It's 9 a.m. to 9 p.m. Six days a week. Six days a week.
什么是996?就是早上9点到晚上9点,每周六天。每周六天。
What is that? 72 hours? Okay. So is this basically the standard in AI companies in Silicon Valley?
那是什么?72小时?好吧。所以这基本上是硅谷AI公司的标准吗?
More and more of this kind of grind mindset? Yeah. I mean, maybe not exactly like that, but I think there is a trend towards it.
越来越多这种苦干心态?是的。我的意思是,也许不完全是那样,但我认为有一种趋势。
And it's interesting. I think it almost flipped because when I was in academia, I felt like that because as a professor, you had to write grants, you had to teach, and you had to do research.
这很有趣。我觉得几乎反过来了,因为当我在学术界时,我有那种感觉,因为作为教授,你必须写拨款申请,必须教学,还必须做研究。
It's like three jobs in one. And it is more than a full-time job if you want to be successful.
就像三份工作合而为一。如果你想成功,这比全职工作还要多。
And I feel like now, like Nathan just said, the professors, in comparison to a lab, I think they have less like even maybe pressure or workload than at a frontier lab
而且我觉得现在,就像Nathan刚才说的,教授们,和实验室相比,我觉得他们的压力甚至工作量可能比前沿实验室要小
because they work a lot they're just so fulfilled
因为他们工作很多,他们就是很满足
but like working with students and having a constant runway of mentorship and like a mission that is very people-oriented
但是像和学生一起工作,有持续的指导机会,以及一个非常以人为本的使命
i think in an era when things are moving very fast and very chaotic it's very rewarding to people
我认为在一个事情发展非常快、非常混乱的时代,这对人们来说是非常有回报的
yeah and i think as a startup i think it's a pressure
是的,而且我认为作为一家初创公司,我觉得这是一种压力
it's like you have to make it and it's like it is really important that people put in the time
就像你必须成功,而且人们投入时间真的很重要
put in the time 常用搭配
投入时间
指为了达成目标而付出必要的时间努力。
but well it is really hard because you have to deliver constantly
但嗯,这真的很难,因为你必须不断地交付成果
and i've been at a startup i had a good time but i don't know if i could do it forever
而且我曾在初创公司工作过,我过得很开心,但我不知道我能不能永远这样做
it's like a interesting pace uh and it's exactly like we talked about in the beginning
这是一种有趣的节奏,呃,这就像我们一开始讨论的那样
these models are leapfrogging each other and they are just constantly like trying to take the next step
这些模型正在互相超越,它们只是不断地试图迈出下一步
compared to the competitors it's just ruthless i think right now
与竞争对手相比,我觉得现在就是很残酷
i think this leapfrogging nature and having multiple players is actually an underrated driver of
我认为这种互相超越的特性以及拥有多个参与者,实际上是一个被低估的驱动因素
Language modeling process where competition is so deeply ingrained to people,
语言建模过程中,竞争在人们心中如此根深蒂固,
and these companies have intentionally created very strong culture, like Anthropic is known to be
而这些公司有意创造了非常浓厚的文化,比如 Anthropic 以
so culturally like deeply committed and organized.
在文化上如此深度投入和有条理而闻名。
I mean, like we hear so little from them, and everybody, it's Anthropic, seems very aligned, and it's like being in a culture that is super tight and having this competitive dynamic
我的意思是,我们很少听到他们的消息,而每个人,Anthropic 似乎非常一致,就像身处一个非常紧密的文化中,拥有这种竞争动态
is like talk about a thing that's going to make you work hard and create things that are better.
就像谈论一件会让你努力工作并创造更好事物的事情。
So I think that this, but that comes at the cost of human capital,
所以我认为这,但代价是人力资本,
comes at the cost of 常用搭配
以……为代价
用于说明获得某样东西的同时会牺牲另一样东西,常用于讨论利弊权衡。
which is like you can only do this for so long, and people are definitely burning out.
也就是说你只能这样做这么久,而人们肯定在倦怠。
burning out 常用搭配
精疲力竭、身心耗尽
形容因长期高强度工作或压力而身心俱疲,常见于职场语境。
I think I've, I wrote a post on burnout as I like, I've tread in and out of this myself, especially trying to like be a manager full mode training, it's a crazy job doing this.
我想我已经,我写了一篇关于倦怠的文章,因为我自己也经历过这些,尤其是试图全职做管理培训,做这个工作很疯狂。
tread in and out of 常用搭配
反复经历、进进出出
形容在某件事或某种状态中反复进出、时有时无,口语中常用来谈个人经历。
The book Apple in China by Patrick McGee, he talked about how hard the Apple engineers worked to set up the supply chains in China.
Patrick McGee 写的《苹果在中国》这本书,他谈到了苹果工程师们为在中国建立供应链付出了多么艰辛的努力。
And he was like, they had saving marriage programs, and he told in a podcast, he was like, people died from this level of working hard.
他说,他们有挽救婚姻的项目,他在一个播客中说,人们因为这种程度的工作努力而死亡。
So I think that it's just like, it's a perfect environment for creating progress based on human expense.
所以我觉得,这就像是一个以人力代价来创造进步的完美环境。
And there's going to be a lot.
而且会有很多。
There's a lot of the human expense is the 996 that we started this with, which is like people do really grind.
很多人力代价就是我们一开始提到的996,就是人们真的在拼命工作。
grind 地道口语
拼命苦干、长时间辛苦工作
口语中形容持续高强度地努力工作,常带辛苦、疲惫的意味。
I also read this book. I think they had a code word for if someone had to go home to spend time with their family to save the marriage.
我也读过这本书。我觉得他们有个暗语,用来指如果有人得回家陪家人以挽救婚姻。
And it's crazy.
这太疯狂了。
Then colleagues understand, OK, this is like red alert for this situation.
然后同事们就明白了,好吧,这种情况就像是红色警报。
red alert 地道口语
红色警报、紧急状态
比喻情况非常紧急、需要立即关注,口语中常用来夸张地表示事态严重。
We have to let that person go home this weekend.
这个周末我们得让那个人回家。
but at the same time I don't think they were forced to work
但与此同时,我不认为他们是被迫工作的
it's really they were so passionate about the product I guess
其实我觉得,他们是真的对产品充满热情
that it is you get into that mindset
以至于你会进入那种心态
get into that mindset 常用搭配
进入那种心态
指逐渐进入某种特定的思维状态或心理模式,常用于描述投入某事时的心理变化。
and I had that sometimes as an academic
我作为学者有时也会有那种心态
but also as an independent person I have that sometimes I overwork
但作为一个独立的人,我有时也会过度工作
and it's unhealthy
这很不健康
I had you know I had back issues I had neck issues
你知道,我有过背部问题,有过颈部问题
because I did not take the breaks that I maybe should have taken
因为我没有休息,也许我本该休息的
but it's not because no one forced me to
但这不是因为有人强迫我
it's because I wanted to work
而是因为我想工作
because it's exciting stuff
因为这是令人兴奋的事情
that's what OpenAI and Anthropics are like they want to do this work
这就是OpenAI和Anthropic的样子,他们想做这项工作
yeah but there's also there's also a feeling a fervor that's building, especially in Silicon Valley, aligned
是的,但也有一种感觉,一种热情正在积聚,尤其是在硅谷,与……一致
with the scaling laws idea where there's this hype where the world will be transformed on a scale of weeks and you want to be at the center of it.
伴随着规模定律的理念,存在这样一种炒作:世界将在几周内被彻底改变,而你想成为这一切的中心。
at the center of it 常用搭配
处于事情的中心
形容身处某个重要事件或趋势的核心位置,常用于表达想参与最前沿的事情。
And then, you know, I have this great fortune of having conversations with a wide variety of human beings.
然后,你知道,我有幸能与各种各样的人交谈。
And from there, I get to see all these bubbles and echo chambers across the world.
从那里,我得以看到世界各地所有这些泡沫和回音室。
echo chambers 常用搭配
回音室(指只听到相同观点的封闭环境)
比喻人们只接触与自己观点一致的信息、听不到不同声音的环境,常用于讨论社交媒体或群体思维。
And it's fascinating to see how we humans form them.
看到我们人类如何形成它们,这非常引人入胜。
And I think it's fair to say that Silicon Valley is a kind of echo chamber, a kind of silo and bubble.
我认为可以说,硅谷是一种回音室,一种筒仓和泡沫。
I think bubbles are actually really useful and effective.
我认为泡沫实际上非常有用且有效。
It's not necessarily a negative thing because it could be ultra productive.
这不一定是一件坏事,因为它可能极其高效。
It could be the Steve Jobs reality distortion field because you just convince each other the breakthroughs are imminent.
它可能是史蒂夫·乔布斯的现实扭曲力场,因为你们只是互相说服对方突破即将到来。
reality distortion field 常用搭配
现实扭曲力场
指某人(尤指有魅力的领导者)能让周围人相信不切实际的想法,源自对乔布斯的描述。
And by convincing each other of that, you make the breakthrough is imminent.
通过互相说服这一点,你使得突破即将到来。
Byron Hobart wrote a book classifying bubbles, but essentially one of them is financial bubbles, which is like speculation, which is bad.
拜伦·霍巴特写了一本关于泡沫分类的书,但基本上其中一种是金融泡沫,就像投机,这是不好的。
And the other one is like, I don't know the term, but effectively for build outs because it pushes people to build these things.
而另一种就像,我不知道术语,但实际上是用于建设扩张,因为它推动人们去建造这些东西。
And I do think AI is in this,
我确实认为AI也参与其中,
but I worry about it transitioning to a financial bubble,
但我担心它会演变成金融泡沫,
which is like... Yeah, but also in the space of ideas, that bubble,
这就像……是的,但在思想领域,那个泡沫,
you are doing a reality distortion field, and that means you are deviating from reality.
你在制造一个现实扭曲力场,这意味着你正在偏离现实。
deviating from reality 常用搭配
偏离现实
指想法或行为与现实脱节,常用于批评脱离实际的判断或群体思维。
And if you go too far from reality, while also working 996,
如果你偏离现实太远,同时还996工作,
and you might miss some fundamental aspects of the human experience, including in Silicon Valley,
你可能会错过人类体验的一些基本方面,包括在硅谷,
and this is a common problem in Silicon Valley, it's a very specific geographic area.
这是硅谷的一个常见问题,它是一个非常特定的地理区域。
You might not understand the Midwest perspective,
你可能不理解中西部视角,
the full experience of all the other different humans in the United States and across the world,
美国及世界各地所有其他不同人类的完整经历,
and you speak a certain way to each other, you convince each other of a certain thing,
你们以某种方式交谈,互相说服某种事情,
And that can get you into real trouble.
这可能会让你陷入真正的麻烦。
get you into real trouble 句型
让你陷入真正的麻烦
get [someone] into real trouble
用于警告某种行为或情况会带来严重后果,语气直接。
Whether AI is a big success and becomes a powerful technology or it's not.
无论AI是巨大成功并成为强大技术,还是并非如此。
In either trajectory, you can get yourself into trouble.
无论哪种轨迹,你都可以让自己陷入麻烦。
So you have to consider all of that.
所以你必须考虑所有这些。
Here you are, a young person trying to decide what you want to do with your life.
你在这里,一个年轻人,试图决定你想用你的生活做什么。
The thing that is, I don't even really understand this,
问题是,我甚至都不太理解这个,
but the SFAI memes have gotten to the point where permanent underclass was one of them,
但SFAI的梗已经发展到永久底层阶级是其中之一,
which was the idea that the last six months of 2025 was the only time to build a durable value in an AI startup or model.
这个梗的意思是,2025年的最后六个月是在AI初创公司或模型中建立持久价值的唯一时机。
Otherwise, all the value will be captured by existing companies and you will therefore be poor,
否则,所有价值都会被现有公司捕获,因此你会很穷,
which that's an example of the SF thing that goes so far.
这就是旧金山那种思维走到极端的例子。
I still think for young people that going to be able to tap into it, if you are really passionate about wanting to have impact in AI,
我仍然认为,对于年轻人来说,如果你真的热衷于在AI领域产生影响,能够接触到它,
being physically in SF is the most likely place for you going to do this,
亲身在旧金山是你最有可能做到这一点的地方,
but it has trade-offs.
但它也有取舍。
trade-offs 常用搭配
取舍、权衡利弊
指为了得到某样东西而必须放弃另一样东西,常用于讨论决策中的利弊。
I think SF is an incredible place.
我认为旧金山是个了不起的地方。
But there is a bit of a bubble.
但有点泡沫。
And if you go into that bubble, which is extremely valuable,
如果你进入那个泡沫,虽然它非常有价值,
just get out also.
也要记得走出来。
Read history books, read literature.
读历史书,读文学。
Visit other places in the world.
去世界上其他地方看看。
Twitter is not, and Substack is not the entire world.
推特不是,Substack也不是整个世界。
I think I would say, one of the people I worked with is moving to SF.
我想我会说,和我共事过的人中有一个要搬去旧金山了。
And it's like, I need to get him a copy of The Season of the Witch,
就像,我得给他一本《女巫的季节》,
which is a history of SF from like 1960 to 1985,
那是一本关于旧金山从1960年到1985年的历史书,
which goes through like the hippie revolution, like they all the um gays kind of taking over the city and that culture emerging,
它讲述了嬉皮士革命,还有那些同性恋者逐渐接管这座城市,以及那种文化的兴起,
and then the HIV AIDS crisis and other things,
然后是艾滋病危机和其他事情,
and it's just like that is so recent and so much turmoil and hurt,
就像,那是如此近的事情,有那么多动荡和伤痛,
but also like love and SF and it's like no one knows about this.
但也有爱和旧金山,就像没人知道这些。
It's a great book, Season of the Witch, which I recommend it.
这是一本好书,《女巫的季节》,我推荐它。
A bunch of my SF friends who do get out recommended it to me.
我一些确实会出门的旧金山朋友向我推荐了它。
And I think that it's just like living there.
我觉得这就像住在那里一样。
Like I lived there and I didn't appreciate this context.
就像我住在那里,却没有体会到这个背景。
And it's just like so recent.
而这就像,如此近的事情。
Yeah.
是的。
Okay.
好的。
Let's, uh, we talked a lot about, we talked a lot about a lot of things, uh, certainly about the things that were exciting last year,
让我们,呃,我们谈了很多,我们谈了很多事情,呃,当然谈了去年令人兴奋的事情,
but this year, uh, one of things you guys mentioned that's exciting is the scaling of text-to-fusion models
但今年,呃,你们提到的一件令人兴奋的事情是文本到融合模型的扩展,
and just a different exploration of text-to-fusion can you talk about what that is and what the possibility holds sort of different kinds of approaches than the current llms
以及只是对文本到融合的不同探索,你能谈谈那是什么以及可能性有哪些,比如不同于当前大型语言模型的不同方法吗?
yeah so we talked a lot about the transformer architecture and the auto-regressive transformer architecture specifically like gpt
是的,所以我们谈了很多关于Transformer架构和自回归Transformer架构,特别是像GPT这样的,
and it doesn't mean no one else is working on anything else so people are always on the let's say look out for the next big thing
这并不意味着没有人在研究其他东西,所以人们总是在,比如说,留意下一个大事件,
look out for the next big thing 常用搭配
留意下一个重大突破或热门事物
在谈论行业趋势或未来机会时,表示持续关注即将出现的重要新事物。
because I think it would be almost like um yeah stupid not to because sure right now the transformer architecture is the thing and it works best and there's right now nothing else out there but you know it's always a good idea to not put all your eggs into one basket
因为我认为如果不这样做几乎就是,呃,是的,愚蠢的,因为当然,现在Transformer架构是主流,它效果最好,目前没有其他东西,但你知道,不要把所有的鸡蛋放在一个篮子里总是一个好主意,
so people are developing other things alternatives to the um autoregressive transformer one of them would be for example text diffusion models
所以人们正在开发其他东西,替代呃自回归Transformer,其中之一例如就是文本扩散模型,
and listeners may know diffusion models from the image generation like stable diffusion popularized it there was
听众可能从图像生成中知道扩散模型,比如Stable Diffusion普及了它,曾经有
like a paper on generating images
像一篇关于生成图像的论文
back then people used GANs, the generative adversarial networks,
当时人们使用GAN,即生成对抗网络,
and then there was this diffusion process where you iteratively denoise an image
然后出现了这种扩散过程,你迭代地对图像去噪
and that resulted in really good quality images over time.
这随着时间的推移产生了非常高质量的图像。
Stable Diffusion was a company,
Stable Diffusion是一家公司,
other companies built their own diffusion models,
其他公司建立了自己的扩散模型,
and then people are now like, okay, can we try this also for text?
然后人们现在就像,好吧,我们能不能也把这个用于文本?
Doesn't, you know, make intuitive sense yet,
你知道,这还不那么直观,
because it feels like, okay, it's not something continuous like a pixel that we can differentiate,
因为它感觉像是,好吧,它不是像像素那样我们可以微分的连续的东西,
it's like a discrete text, so how do we implement that denoising process?
它像是离散的文本,那么我们如何实现那个去噪过程呢?
But it's kind of like similar to, uh, the BERT models by Google,
但它有点像,呃,谷歌的BERT模型,
like when you go back to the original Transformer,
就像当你回到原始的Transformer,
and so they were like, the encoder and the decoder, the decoder
所以他们就像,编码器和解码器,解码器
is what we are using right now in GPT and so forth.
是我们现在在GPT等中使用的。
The encoder, it's more like, um, a parallel, let's say, technique
编码器,它更像,嗯,一种并行的,比如说,技术
where you have multiple tokens that you fill in in parallel
你有多个标记,你并行地填充它们
instead. So GPT models, they do auto-regressive, one token at a time,
相反。所以GPT模型,它们做自回归,一次一个标记,
you complete the sentence one token at a time,
你一次一个标记地完成句子,
and in BERT models, you have a text that's a sentence that has gaps
而在BERT模型中,你有一个文本,是一个有缺口的句子
You like mask them out, and then one iteration is filling in these gaps,
你像是把它们遮住,然后一次迭代就是填补这些空缺,
and text diffusion is kind of like that, where you are starting with
而文本扩散就有点像那样,你从
let's say some random text, and then you are filling in the missing parts,
比如说一些随机文本开始,然后你填补缺失的部分,
or you're refining them iteratively, and you have multiple iterations,
或者你迭代地精炼它们,你有多次迭代,
and the cool thing here is that this can do multiple tokens at the same time,
而这里很酷的一点是,它可以同时处理多个词元,
so it's kind of like the promise of having it more efficient now.
所以这有点像现在让它更高效的承诺。
The trade-off is, of course, well, how good is the quality? It might be faster,
当然,权衡在于,质量有多好?它可能更快,
and then now you have this dimension of the denoising process:
然后现在你有了去噪过程的这个维度:
the more steps you do, the better the text becomes,
你做的步骤越多,文本就变得越好,
the more steps you do, the better the text becomes 句型
步骤越多,文本就变得越好
the more [something] you do, the better [something] becomes
用于表达两个变量同向变化,即一个越多,另一个也越怎样。
um, and people, you know, I mean, you can scale in different ways.
嗯,而且人们,你知道,我的意思是,你可以用不同的方式扩展。
They try to see if that is maybe a valid alternative to the autoregressive model
他们试图看看那是否可能是自回归模型的一个有效替代方案,
in terms of giving you the same quality for less compute.
在给你相同质量但计算量更少方面。
Right now, I think it's, you know, there are papers that suggest,
现在,我认为,你知道,有论文表明,
okay, if you want to get the same quality,
好吧,如果你想获得相同的质量,
you have to crank up the denoising steps, and then you end up spending the same compute you would spend on an autoregressive model.
你必须增加去噪步骤,然后你最终花费的计算量会和自回归模型一样。
The other downside is, well, it's parallel, which sounds appealing,
另一个缺点是,嗯,它是并行的,听起来很吸引人,
but some tasks are not parallel, like, you know, like reasoning tasks, tool use, maybe where you have to ask a code interpreter to give you an intermediate result,
但有些任务不是并行的,比如,你知道,像推理任务、工具使用,也许是你必须让代码解释器给你一个中间结果,
and that is kind of tricky with diffusion models, so there are some hybrids, but the main idea is can we parallelize it and so interesting avenue I think right now
而这在扩散模型上有点棘手,所以有一些混合模型,但主要想法是,我们能不能把它并行化,所以我认为现在这是一个有趣的途径
there are mostly research, uh, let's say models out there like Lada and some other ones I saw some by startup, some deployed models there is no big uh diffusion model at scale yet
目前大多还是研究性的,呃,比如说像Lada这样的模型,还有一些其他的,我看到一些创业公司的,一些部署的模型,还没有大规模的呃扩散模型
like you know like Gemini, ChatGPT scale in that level, but there was an announcement by Google or like a site where they said they are launching Gemini diffusion and they put it into context of their I think Nano 2 model and then they said basically for the same quality on most benchmarks we can generate things much faster so
比如你知道像Gemini、ChatGPT那种规模,但谷歌有一个公告,或者说一个网站,他们说他们要推出Gemini扩散模型,并把它放在他们的Nano 2模型的背景下,然后他们基本上说,在大多数基准测试上,为了同样的质量,我们可以更快地生成东西,所以
you mentioned what's next I don't think the text diffusion model is going to replace autoregressive LMs
你提到了下一步,我不认为文本扩散模型会取代自回归语言模型
but it will be something maybe for quick cheap at scale tasks maybe the free tier in future will be something like that I think there's
但它可能会用于快速、廉价的大规模任务,也许未来的免费层会是那样的,我想有
a couple examples where it's, I've heard that it's actually been started to be used.
有几个例子,我听说它实际上已经开始被使用了。
I think to paint an example of why this is so much better,
我想举个例子说明为什么这个要好得多,
for example, when GPT-5 is taking 30 minutes to respond, it's generating one token at a time.
比如说,当GPT-5需要30分钟才能回应时,它是一次生成一个token。
And this diffusion idea is essentially generate all of those tokens in the completion in one batch, which is why it could be way faster.
而这种扩散的思路基本上就是一次性批量生成完成内容中的所有token,这就是为什么它可能会快得多。
And I think it could be suited.
我觉得它可能会很适合。
The startups I'm hearing are like code startups where you have a code base and you have somebody that's effectively vibe coding,
我听说的一些初创公司是像代码类的初创公司,你有一个代码库,然后有个人基本上就是在凭感觉写代码,
and they say, make this change.
然后他们说,做这个改动。
And a code diff is essentially a huge reply from the model, but it doesn't have to have that much external context.
而代码差异本质上就是模型给出的一个巨大的回复,但它不需要那么多外部上下文。
And you can get it really fast by using these diffusion models.
而通过使用这些扩散模型,你可以非常快地得到它。
So that's what I've heard of, one example is that they use these text diffusion to generate really long diffs
所以这就是我听说的情况,一个例子是他们用这些文本扩散来生成很长的差异,
because doing it with a autoregressive model would take minutes.
因为用自回归模型来做的话会花上几分钟。
And that time for like a user facing product causes a lot of churn.
而那样的时间对于面向用户的产品来说会导致大量用户流失。
So like every second you lose a lot of users.
所以就像每一秒你都会流失很多用户。
So I think that's going to be this thing where it's going to grow and have some applications.
所以我觉得这将会是那种会发展起来并有一些应用的东西。
But I actually thought that different types of models were going to be used for different things sooner than they have been.
但我其实以为不同类型的模型会更早地被用于不同的事情,比它们实际发生的时间要早。
So I kind of trade off.
所以我有点权衡取舍。
I think that the tool use point is the one that's stopping them from being like most general purpose
我认为工具使用的关键点正是阻止它们成为最通用目的的那个因素
because like Cloud Code and this Hatchipute with Search, like the autoregressive chain is interrupted with some external tool.
因为像 Cloud Code 和这个带搜索的 Hatchipute,自回归链被一些外部工具打断了。
And I don't know how to do that with the diffusion setup.
而我不知道如何在扩散设置中做到这一点。
So what's the future of tool use this year and in the coming years?
那么今年以及未来几年,工具使用的未来是什么?
Do you think there's going to be a lot of developments there, how that's integrated to the entire stack?
你认为那里会有很多发展吗,以及它如何整合到整个技术栈中?
I do think right now, I mean, it's mostly on the proprietary LLM side,
我确实认为现在,我的意思是,它主要是在专有 LLM 方面,
but I think we will see more of that in the open source tooling.
但我认为我们会在开源工具中看到更多这样的情况。
And I think, I mean, it is a huge unlock
而且我认为,我的意思是,这是一个巨大的解锁
because then you can really outsource certain tasks from just memorization to actual, you know,
因为那样你就可以真正将某些任务从单纯的记忆外包到实际的,你知道,
like instead of having the LLM memorized, what is 23 plus 5? Just use a calculator.
就像与其让 LLM 记住,23 加 5 等于多少?直接用计算器。
So do you think that can help solve hallucination?
那么你认为这能帮助解决幻觉吗?
Not solve it, but reduce it.
不是解决它,而是减少它。
So still the LLM needs to know when to ask for a tool call.
所以 LLM 仍然需要知道何时请求工具调用。
And the second one is, well, it doesn't mean the internet is always correct.
第二个是,嗯,这并不意味着互联网总是正确的。
You can do a web search, but let's say I asked who won the World Cup in, let's say, 1998.
你可以做网络搜索,但假设我问谁赢了1998年的世界杯。
let's say 地道口语
比如说,假设
口语中用来举例或提出假设情况,语气随意。
It still needs to find the right website and get the right information.
它仍然需要找到正确的网站并获取正确的信息。
So you can still go to the incorrect website and give me incorrect information.
所以你仍然可能访问错误的网站并给我错误的信息。
So I don't think it will fully solve that, but it is improving it in that sense.
所以我认为它不会完全解决这个问题,但在这个意义上它正在改进。
in that sense 常用搭配
在那个意义上,从那个角度来说
用于限定或澄清前面说法的适用范围。
And so another cool paper earlier this year, I think it was December 31st.
所以今年早些时候的另一篇很酷的论文,我想是12月31日。
So it's not technically 2026, but close.
所以严格来说不是2026年,但很接近。
So like the recursive language model.
所以就像递归语言模型。
That's a cool idea to kind of take this even a bit further.
这是一个很酷的想法,可以把这个再推进一步。
take this even a bit further 常用搭配
把这个再推进一步
表示在现有想法基础上进一步拓展或深化。
So just to explain, so Nathan, you also mentioned earlier, it's harder to do cool research in academia because of the compute budget.
所以解释一下,Nathan,你之前也提到,由于计算预算,在学术界做酷研究更难。
if i recall correctly they did everything with gpd5 so they didn't even use local models but the idea is let's say if a long context task instead of having the llm solve all of it in like one shot or even like in a chain you break it down into subtasks you have the llm decide when like what is a good let's say subtask and then recursively call an llm to solve that and i think something like that also then adding tools and you know each one maybe you have like a huge q and
如果我没记错的话,他们用gpd5做了一切,所以甚至没有使用本地模型,但想法是,假设一个长上下文任务,不是让LLM一次性解决所有问题,甚至像链式那样,你把它分解成子任务,让LLM决定什么时候,比如什么是好的子任务,然后递归调用LLM来解决它,我认为类似这样的东西,然后添加工具,你知道每个可能你有一个巨大的q和
if i recall correctly 常用搭配
如果我没记错的话
在陈述记忆中的信息时,表示不确定或礼貌地留有余地。
break it down into subtasks 常用搭配
把它分解成子任务
用于描述把复杂任务拆分成更小、更易处理的部分。
a task. So each one goes to the web and gathers information and then you pull it at the end together and stitch it back together.
一个任务。所以每一个都会上网收集信息,然后你在最后把它拉取到一起,再拼接回去。
stitch it back together 常用搭配
把它重新拼接起来
比喻把分散的部分重新组合成一个整体。
Like where I think there's going to be a lot of unlock using things like that where you don't necessarily improve the LLM itself.
就像我认为会有很多解锁,使用类似那样的东西,你不一定改进LLM本身。
You improve how the LLM is used and what the LLM can use.
你改进的是LLM的使用方式以及LLM能使用什么。
One downside right now with tool use is you have to give the LLM permission to use tools.
目前工具使用的一个缺点是,你必须给LLM使用工具的权限。
And that will take some trust, especially if you want to unlock things like having an LLM answer emails for you.
而这需要一些信任,特别是如果你想解锁像让LLM替你回复邮件这样的事情。
Not even answer, but just sort them for you or select them for you or something like that.
甚至不是回复,而只是帮你分类,或者帮你筛选,或者类似的事情。
I don't know if I would today give an LLM access to my emails, right?
我不知道今天我会不会给LLM访问我邮件的权限,对吧?
I mean, this is like a huge risk.
我的意思是,这就像是一个巨大的风险。
I think there's a cool one last point on the tool use thing.
我认为关于工具使用这件事,还有最后一个很酷的点。
I think that you hinted at this and we've both come at this in our own ways is that the open versus closed models use tools in very different ways
我认为你暗示了这一点,而我们俩都以自己的方式探讨过,那就是开源模型和闭源模型使用工具的方式非常不同。
where open models, people go to Hugging Face and you download the model and then the person's going to be like, oh, what tool do I want?
在开源模型中,人们去Hugging Face下载模型,然后这个人就会想,哦,我想要什么工具?
And I don't know, Exa is my preferred search provider,
我不知道,Exa 是我首选的搜索提供商,
but somebody else might care for a different search startup
但其他人可能喜欢不同的搜索初创公司
where you release a model.
在那里你发布一个模型。
It needs to be useful for multiple tools for multiple use cases,
它需要对多种工具、多种用例都有用,
which is really hard because you're making like a general reasoning engine model,
这真的很难,因为你在做一个通用的推理引擎模型,
which is actually what GPT-OSS is good for.
而这正是 GPT-OSS 所擅长的。
But on the closed models, you're deeply integrating the specific tool into your experience.
但在闭源模型上,你要将特定工具深度整合到你的体验中。
And I think that open models will struggle to replicate some of the things that I like to do with closed models,
我认为开源模型将难以复制我喜欢用闭源模型做的一些事情,
which will be like, I don't know, you can reference a mix of public and private information
比如,我不知道,你可以引用公共和私人信息的混合
and something that I keep trying every three to six months
以及我每三到六个月都会尝试的事情
that I try like codecs on the web,
比如我在网上尝试编解码器,
which is just prompting a model to make an update to some GitHub repository that I have.
这只是提示模型对我拥有的某个 GitHub 仓库进行更新。
And it's just like that set of secure cloud environment
而这就像那套安全的云环境
is just so nice for just like, send it off and do this thing and then come back to me.
真的很好,就像,把它发出去,做这件事,然后回来找我。
send it off 常用搭配
把它发出去
指把任务或信息发送出去让别人处理,自己不用盯着。
And these will probably help define some of the local open and closed niches,
这些可能会帮助定义一些本地的开放和封闭利基市场,
but I think initially, because there was such a rush to get these tool use working,
但我认为最初,因为大家急于让这些工具使用功能运作起来,
that the open models were on the back foot, which is kind of inevitable.
所以开放模型处于劣势,这有点不可避免。
on the back foot 地道口语
处于劣势,被动
形容在竞争或局势中处于不利、防守的位置。
I think there's so much research, so many resources in these frontier labs,
我认为这些前沿实验室有如此多的研究、如此多的资源,
but will be fun when the open models solve this,
但当开放模型解决这个问题时将会很有趣,
because it's going to necessitate like a bit more flexible and potentially interesting model that might work with this recursive idea to like be an orchestrator and a tool used model.
因为它将需要一种更灵活、可能更有趣的模型,这种模型可能利用这种递归思想来充当编排器和工具使用的模型。
So hopefully the necessity drives some interesting innovation there.
所以希望这种必要性能够推动一些有趣的创新。
So continual learning.
那么,持续学习。
This is a longstanding process. topic, important problem, I think that increases in importance as the cost of training of the models goes up.
这是一个长期的过程。话题,重要问题,我认为随着模型训练成本的上升,其重要性也在增加。
So can you explain what continual learning is and how important it might be this year and in the coming years to make progress?
那么你能解释一下什么是持续学习,以及它在今年和未来几年对于取得进展可能有多重要吗?
This relates a lot to this kind of SF zeitgeist
这与这种科幻时代精神有很大关系,
of what is AGI, which is artificial general intelligence,
即什么是AGI,即人工通用智能,
and what is ASI, artificial superintelligence,
以及什么是ASI,人工超级智能,
and what are the language models that we have today capable of doing?
以及我们今天拥有的语言模型能够做什么?
I think the language models can solve a lot of tasks,
我认为语言模型可以解决很多任务,
but a key milestone among the AI community is essentially when AI could replace any remote worker taking in information and solving digital tasks and doing them.
但AI社区的一个关键里程碑基本上是当AI能够取代任何远程工作者,接收信息、解决数字任务并完成它们。
And the limitation that's highlighted by people is that a language model will not learn from feedback the same way that an employee is.
而人们强调的局限是,语言模型不会像员工那样从反馈中学习。
So if you hire an editor, the editor will mess up, but you will tell them.
所以如果你雇一个编辑,编辑会搞砸,但你会告诉他们。
And if you hired a good editor, they don't do it again.
如果你雇了一个好编辑,他们不会再犯。
But language models don't have this ability to modify themselves and learn very quickly.
但语言模型没有这种能力来修改自己并快速学习。
So the idea is if we're going to actually get to something that is a true general adaptable intelligence that can go into any remote work scenario, it needs to be able to learn quickly from feedback and on-job learning.
所以想法是,如果我们要真正达到一种真正的通用适应性智能,能够进入任何远程工作场景,它需要能够从反馈和在职学习中快速学习。
I'm personally more bullish on language models by being able to just provide them with very good context.
我个人更看好语言模型,因为能够直接为它们提供非常好的上下文。
bullish on 常用搭配
看好,对……持乐观态度
常用于商业或投资语境,表示对某事物前景有信心。
You said, like you maybe offline said that you can write extensive documents to models
你说过,就像你也许离线时说的,你可以给模型写详尽的文档,
where you say, I have all this information.
你在里面说,我有所有这些信息。
Here's all the blog posts I've ever written.
这是我写过的所有博客文章。
I like this type of writing.
我喜欢这种写作风格。
My voice is based on this.
我的声音就是基于这个。
But a lot of people don't provide this to models,
但很多人不向模型提供这些,
and the models weren't designed to take this amount of context previously.
而且这些模型以前并不是为了接受这么大量的上下文而设计的。
Like the agentic models are just starting.
就像智能体模型才刚刚起步。
So it's this kind of trade-off of, do we need to update the weights of this model with this continual learning thing to make them learn fast?
所以这是一种权衡:我们是否需要通过这种持续学习来更新这个模型的权重,让它们学得快?
Or the counter argument is, we just need to provide them with more context and information and they will have the appearance of learning fast by just having a lot of context and being very smart.
或者反方观点是,我们只需要给它们提供更多的上下文和信息,它们就会因为拥有大量上下文并且非常聪明而显得学得很快。
So we should mention the terminology here.
所以我们应该在这里提一下术语。
So continual learning refers to changing the weights continuously so that the model adapts, adjusts based on the new incoming information,
所以持续学习指的是不断改变权重,使模型根据新输入的信息进行适应和调整,
does so continually and rapidly and frequently and so on.
持续地、快速地、频繁地这样做,等等。
And then the thing you mentioned on the other side of it is generally be referred to as in-context learning.
然后你提到的另一面通常被称为上下文学习。
As you learn stuff, there's a huge context window.
当你学习东西时,有一个巨大的上下文窗口。
You can just keep loading it with extra information every time you prompt the system,
你可以在每次提示系统时不断加载额外信息,
which I think both are legitimately can be seen as learning.
我认为两者都可以合理地被视为学习。
It's just a different place where you're doing the learning.
只是进行学习的地方不同而已。
I think, to be honest with you, continual learning, the updating of weights,
老实说,我认为持续学习,也就是更新权重,
to be honest with you 地道口语
老实说,跟你说实话
口语中用于坦诚表达个人看法,常放在句首或句中。
we already have that in different flavors.
我们已经以不同的形式拥有了它。
I think the distinction here is, do you do that on a personalized custom model for each person, or do it on a global model scale,
我认为这里的区别在于,你是在为每个人做个性化的定制模型,还是在全球模型的规模上做,
on a global model scale 常用搭配
在全球模型的规模上
用于对比个人层面与全球/整体层面的做法,常见于技术或商业讨论。
and I think we have that already with going from GPT-5 to 5.1 and 5.2.
而我认为我们从GPT-5到5.1再到5.2,已经有了这一点。
It's maybe not immediate, but it is like a curated update, a quick curated update,
这可能不是即时的,但它像是一种精心策划的更新,一次快速的精心策划的更新,
where there was feedback by the things that couldn't do, feedback by the community, they updated the weights, next model, and so forth.
其中有来自做不到的事情的反馈,来自社区的反馈,他们更新了权重,下一个模型,等等。
and so forth 常用搭配
等等,诸如此类
列举若干例子后表示还有更多类似情况,口语和书面均可。
So it is, I mean, kind of like a flavor of that,
所以它其实就是,我的意思是,有点像那种形式,
kind of like 地道口语
有点像,差不多是
口语中用来弱化语气、近似描述某事物。
um, other even finer grade example, a finer grained example is like RLVR.
嗯,另一个更细粒度的例子,一个更细粒度的例子就像RLVR。
You run it, it updates the problem is you can't just do that for each person,
你运行它,它更新,问题是你不能只为每个人这样做,
because it would be too expensive to update the weights for each person.
因为为每个人更新权重的成本太高了。
And I think that's the problem. So unless you get, I mean, even at OpenAI scale, building the data centers,
而我认为这就是问题所在。所以除非你能,我的意思是,即使在OpenAI的规模上,建设数据中心,
it would be too expensive. I think that is only feasible once you have something on the device,
那也会太贵了。我认为只有当你在设备上有东西时,这才可行,
where the cost is on the consumer, like what Apple tried to do with the Apple Foundation models, putting them on the phone,
成本由消费者承担,就像苹果试图用Apple Foundation模型做的那样,把它们放在手机上,
and then they learn from the experience.
然后它们从经验中学习。
A bit of a related topic, but this kind of maybe anthropomorphized term, but memory.
有点相关的话题,但这种可能被拟人化的说法,但记忆。
What are different ideas of the mechanism of how to add memory to these systems as you're increasingly seeing so?
随着你越来越多地看到这种情况,关于如何给这些系统添加记忆的机制,有哪些不同的想法?
So personalized memory especially.
尤其是个人化记忆。
So right now it's mostly like context, basically stuffing things into the context and then just recalling that.
所以现在基本上就是上下文,基本上就是把东西塞进上下文里,然后只是回忆它。
But again, I think, well, it's expensive because you have to, I mean, you can cash it, but still you spend tokens on that.
但话说回来,我觉得,嗯,这很贵,因为你必须,我是说,你可以缓存它,但你仍然要为此花费token。
And the second one is you can only do so much.
第二点是你能做的有限。
you can only do so much 句型
你能做的有限
[someone] can only do so much
表示能力或效果有上限,无法无限做下去。
I think it's more like a preference or a style.
我觉得这更像是一种偏好或风格。
I mean, a lot of people do that when they solve math problems.
我是说,很多人在解数学题时就是这么做的。
You say, it's basically you can add previous knowledge and stuff, but you also give it certain preference prompts, do what I preferred last time, whatever, like something like that.
你说,基本上就是你可以添加先前的知识和东西,但你也会给它某些偏好提示,做我上次偏好的那样,随便什么,类似那样的。
But it doesn't unlock new capabilities.
但它不会解锁新的能力。
So for that, one thing people do use still is LoRa, LoRa adapters.
所以为此,人们仍然使用的一个东西是LoRa,LoRa适配器。
These are basically, instead of updating the whole weight matrix, there are two smaller weight matrices that you kind of have in parallel or overlay.
这些基本上就是,不是更新整个权重矩阵,而是有两个较小的权重矩阵,你有点像是并行或叠加地拥有它们。
It's like the delta.
就像增量一样。
But yeah, you can do that to some extent.
但是的,你可以在某种程度上做到这一点。
to some extent 常用搭配
在某种程度上
表示部分成立、并非完全如此,常用于缓和陈述。
But then again, it is economics.
但话说回来,这还是经济学问题。
So there were also papers, for example, LoRa learns less but forgets less.
所以也有一些论文,比如LoRa学得少但忘得也少。
It's like, you know, it's no free lunch.
就像,你知道,天下没有免费的午餐。
If you want to learn more, you need to use more weights, but it gets more expensive.
如果你想学得更多,你就需要用更多的权重,但那样会变得更贵。
And then again, if you learn more, you forget more.
而且反过来说,如果你学得更多,你忘得也更多。
And it's like you have to find that Goldilocks zone, basically.
而且就像你必须找到那个刚刚好的区域,基本上是这样。
We haven't really mentioned it much, but implied in this discussion is context length also.
我们其实没怎么提到它,但在这个讨论中隐含的还有上下文长度。
Is there a lot of innovations that's possible there?
那里有可能有很多创新吗?
I think the colloquially accepted thing is that it's a compute and data problem
我认为大家普遍接受的说法是,这是一个计算和数据的问题,
where you can, and sometimes like small architecture things, which are like attention variance.
你可以,有时候像一些小架构上的东西,比如注意力方差。
So if you have, we talked about like hybrid attention models, which is essentially
所以如果你有,我们谈过像混合注意力模型,这本质上就是
if you have what looks like a state space model within your transformer.
如果你的transformer里有一个看起来像状态空间模型的东西。
And like those are better suited because you have to spend less compute to model the furthest along token.
而像那些更适合,因为你必须花更少的计算来建模最远的那个token。
And I think that, but those aren't free because they have to be accompanied by a lot of compute or the right data.
我认为,但那些并不是免费的,因为它们必须伴随着大量的计算或合适的数据。
So how many sequences of 100,000 tokens do you have in the world and where do you get these?
那么世界上有多少个10万token的序列,你从哪里得到这些呢?
And I think it just ends up being pretty expensive to scale them.
而我认为,扩展它们最终会变得相当昂贵。
ends up being 常用搭配
最终变成,结果是
描述经过一系列过程后的最终结果,口语常用。
So we've gotten pretty quickly to a million tokens of input context length.
所以我们很快就达到了一百万个token的输入上下文长度。
And I would expect it to keep increasing and get to 2 million or 5 million this year.
我预计它会继续增长,今年达到两百万或五百万。
But I don't expect it to go to 100 million.
但我不认为它会达到一亿。
That would be a true breakthrough.
那将是一个真正的突破。
And I think those breakthroughs are possible.
我认为那些突破是可能的。
The continual learning thing, I think of as a research problem
持续学习这件事,我认为是一个研究问题,
where there could be a breakthrough that just makes transformers work way better at this.
可能会有一种突破,让transformers在这方面表现得好得多。
And it's cheap.
而且它很便宜。
These things could happen with so much scientific attention,
这些事情在如此多的科学关注下可能会发生,
but turning the crank, it'll be consistent increases over time.
但持续努力的话,它会随着时间稳定增长。
I think also looking at the extremes, I think there's again no free lunch.
我认为,从极端情况来看,同样没有免费的午餐。
So the one extreme to make it cheap, you have, let's say, an RNN that has a single state
所以一个极端是让它便宜,比如说,你有一个只有单一状态的RNN,
where you save everything from the previous stuff.
你把之前所有的东西都保存在里面。
It's like a specific fixed size thing.
它就像一个特定固定大小的东西。
So you never really grow the memory because you are stuffing everything into one state.
所以你永远不会真正增长记忆,因为你把所有东西都塞进一个状态里。
But then the longer the context gets, the more information you forget
但上下文越长,你忘记的信息就越多,
the longer the context gets, the more information you forget 句型
上下文越长,你忘记的信息就越多
the [comparative] [subject] gets, the more [something] [happens]
用于表达两个变化成正比:越……就越……。
because you can't compress everything into one state.
因为你无法把所有东西都压缩进一个状态里。
Then on the other hand, you have the transformers, which try to remember every token,
另一方面,你有transformers,它们试图记住每一个token,
which is great sometimes if you want to look up specific information,
有时如果你想查找特定信息,这很好,
but very expensive because you have the KV cache that grows,
但非常昂贵,因为你的KV缓存会增长,
the dot product that grows.
点积也会增长。
But then, yeah, like you said, the Mamba layers,
但然后,是的,就像你说的,Mamba 层,
I mean, they kind of have the same problem,
我的意思是,它们有点有同样的问题,
I would say, like an RNN, you try to compress everything into one state,
我会说,像 RNN 一样,你试图把所有东西压缩到一个状态里,
you're a bit more selective there.
你在那里更有选择性一些。
But then I think it's like this Goldilocks zone again with Unimotron 3,
但然后我认为这又像是 Unimotron 3 的 Goldilocks 区域,
they found like a good ratio of how many attention layers do you need for the global information
他们找到了一个很好的比例,关于你需要多少注意力层来获取全局信息
where everything is accessible compared to having these compressed states
在那里一切都可以访问,相比于拥有这些压缩状态
and I think that's how I think we will scale more by finding better let's say ratios in Goldilocks zone
而且我认为这就是我认为我们将通过找到更好的,比如说,Goldilocks 区域中的比例来扩展更多
like between um like computing uh making it cheap enough to run
比如在嗯比如计算呃让它足够便宜以运行之间
but then also making it powerful enough to be useful
但然后也要让它足够强大以有用
and one more plug here um the recursive language model paper
还有一个补充,嗯,递归语言模型论文
that is one of the papers that tries to kind of address the long context thing
那是试图解决长上下文问题的论文之一
so what they found is essentially instead of stuffing everything into this long context
所以他们发现的基本上是,与其把所有东西塞进这个长上下文
if you break it up into these smaller multiple smaller tasks
如果你把它分解成这些更小的多个更小的任务
so you save memory by having multiple smaller calls
所以你通过有多个更小的调用来节省内存
you can get actually better accuracy than having the LM try everything all at once
你实际上可以获得比让语言模型一次性尝试所有东西更好的准确性
all at once 常用搭配
一次性,同时
表示同时处理或发生所有事情,而非分步进行。
I mean it's a new paradigm
我的意思是这是一个新范式
we will see you know there might be other flavors of that
我们会看到,你知道,可能会有其他变体
flavors of 常用搭配
……的变体或不同版本
口语中表示某事物的不同形式或类型,常用于技术讨论。
So I think with that we will still make improvement on long context,
所以我认为,有了那个,我们仍然会在长上下文上取得改进,
make improvement on 常用搭配
在……方面取得改进
用于描述在某个领域或方面取得进步,较正式。
but then also like Nathan said,
但然后也像Nathan说的,
I think the problem is for pre-training itself:
我认为问题在于预训练本身:
we don't have as many long context documents
我们没有那么多长上下文文档,
as other documents,
像其他文档那样,
so it's harder to study
所以更难研究
basically, um, how LMs behave and stuff like that on, on that level,
基本上,嗯,语言模型在那个层面上如何表现之类的,
and stuff like that 地道口语
以及诸如此类的东西
口语中用来列举类似事物,表示不一一详述。
like there are some rules of thumb
比如有一些经验法则,
rules of thumb 常用搭配
经验法则
指基于实践而非精确理论的一般性指导原则。
where essentially you pre-train a language model
基本上你预训练一个语言模型,
like, oh no, we pre-trained like 8k context length,
比如,哦不,我们预训练了8k上下文长度,
and then extended to 32k with training,
然后通过训练扩展到32k,
and there's some rules of thumb where you're just like essentially doubling the training context length
而且有一些经验法则,基本上你只是将训练上下文长度加倍,
takes like 2x compute,
需要大约2倍计算量,
and then you can normally like 2 to 4x the context length again,
然后你通常可以再次将上下文长度增加2到4倍,
so I think a lot of it ends up being kind of compute bound
所以我认为很多最终都是计算量受限的,
ends up being 常用搭配
最终成为
表示经过一系列过程后最终的结果或状态。
at pre-training, which is in this like we talked about this,
在预训练时,也就是在这个,就像我们谈过的,
everyone talks about this big increase in compute for the top labs this year,
每个人都在谈论今年顶级实验室计算量的大幅增加,
and that should reflect in some longer context windows,
而这应该会反映在一些更长的上下文窗口中,
but I think on the post-training side,
但我认为在后训练方面,
there's some more interesting things, which is as we have agents,
有一些更有趣的事情,那就是当我们有智能体时,
the agents are going to manage this context on their own,
智能体将自己管理这个上下文,
on their own 常用搭配
自行,独立地
表示不需要外部帮助,自己完成某事。
where now people that use cloud code
现在使用Cloud Code的人
a lot dread the compaction, which is when Claude takes its entire full 100,000 tokens of work and compacts it into bulleted list.
很多人害怕压缩,就是当Claude把它全部10万个token的工作压缩成项目符号列表的时候。
But what the next models will do, I'm just not a novel, I'm sure people are already working on this, is essentially the model can control when it compacts and how.
但下一代模型会怎么做,我并不是第一个想到的,我相信已经有人在研究了,本质上就是模型可以控制何时以及如何压缩。
So you can essentially like train your RL algorithm where compaction is an action where it shortens the history, and then the problem formulation will be I want to keep the maximum evaluation scores that I have gotten while the model compacts its history to the minimum length.
所以你可以基本上训练你的RL算法,把压缩作为一个动作来缩短历史,然后问题表述就是:我希望在模型将历史压缩到最小长度的同时,保持我获得的最大评估分数。
Because then you have the minimum amount of tokens that you need to do this kind of compounding auto-regressive prediction.
因为这样你就有了进行这种复合自回归预测所需的最少token数量。
So there's actually pretty nice problem setups in this where these agentic models learn to use their context in a different way than just plow forward.
所以这里其实有很好的问题设定,这些智能体模型学会以不同于单纯向前推进的方式使用它们的上下文。
plow forward 常用搭配
一味向前推进
比喻不顾情况地持续前进,不调整策略。
One interesting also recent example would be DeepSeq version 3.2, where they had like the sparse attention mechanism, where they have essentially like a very efficient, small, lightweight indexer.
另一个有趣的近期例子是DeepSeq 3.2版本,他们采用了稀疏注意力机制,本质上有一个非常高效、小巧、轻量级的索引器。
And instead of attending to all the tokens, it selects, okay, what tokens do I actually need?
它不是关注所有token,而是选择,好吧,我实际需要哪些token?
I mean, it almost comes back to the original idea of attention where you are selective,
我的意思是,它几乎又回到了注意力最初的理念,即你是有选择性的,
but attention is always on. You have maybe zero weight on some of them, but you use them all.
但注意力始终是开启的。有些可能权重为零,但你全都在用。
But they are even more like, okay, let's just mask that out or not even do that.
但它们更像是,好吧,我们干脆把它屏蔽掉,或者甚至不这么做。
And even with sliding window attention, that is also kind of like that idea.
而即使是滑动窗口注意力,那也差不多是同样的理念。
You have that rolling window where you keep it fixed because you don't need everything all the time.
你有那个滚动的窗口,你让它保持固定,因为你不需要一直拥有所有东西。
Occasionally, some layers you might, but it's wasteful.
偶尔,某些层你可能会,但那是浪费的。
But right now, I think, yeah, if you use everything, you're on the safe side.
但就目前而言,我觉得,是的,如果你全都用上,你就稳妥了。
on the safe side 常用搭配
稳妥起见
表示为了保险而采取更保守的做法。
It gives you the best bang for the buck because you never miss information.
它能让你获得最高的性价比,因为你永远不会错过信息。
And right now, I think this year will be more also the year figuring out, like you said, how to be more smart about that.
而现在,我觉得今年也更将是摸索的一年,就像你说的,如何在这方面变得更聪明。
I think right now, people want to have the next state of the art.
我觉得现在,人们想要拥有下一个最先进的东西。
And the state of the art happens to be the brute force expensive thing.
而最先进的东西恰好就是那种蛮力、昂贵的东西。
And then once you have that, like you said, keep that accuracy,
然后一旦你有了那个,就像你说的,保持那种准确性,
but let's see how we can do that cheaper now, like tricks, you know.
但让我们看看现在怎么能更便宜地做到这一点,比如一些技巧,你知道的。
Yeah, all this scaling thing.
是的,所有这些扩展的事情。
Like the reason we get the Cod 4.5 Sonnet model first is because you can train it faster
就像我们首先得到 Cod 4.5 Sonnet 模型的原因是因为你可以更快地训练它
and you're not hitting these compute walls as soon and they can just try a lot more things and get the model faster, even though the bigger model is actually better.
而且你不会那么快碰到这些计算瓶颈,他们可以尝试更多东西,让模型更快,尽管更大的模型实际上更好。
hitting these compute walls 常用搭配
碰到计算瓶颈
比喻遇到计算资源的限制,无法继续扩展。
I think we should say that there's a lot of exciting stuff going on in the AI space.
我认为我们应该说,AI 领域正在发生很多令人兴奋的事情。
My mind has recently been really focused on robotics.
我最近的心思真的都集中在机器人技术上。
So today really almost entirely didn't talk about robotics.
所以今天几乎完全没有谈到机器人技术。
There's a lot of stuff on image gen, video generation.
有很多关于图像生成、视频生成的内容。
I think it's fair to say that the most exciting research work in terms of the amount, intensity, fervor is in the LLM space,
我认为可以公平地说,就数量、强度和热情而言,最令人兴奋的研究工作是在 LLM 领域,
it's fair to say 句型
可以公平地说
it's fair to say that [clause]
用于表达一个客观合理的判断,常引导一个普遍认可的观点。
which is why I think it's justified for us to really focus on the LLM that we're discussing.
这就是为什么我认为我们真正专注于我们正在讨论的 LLM 是合理的。
But it would be nice to bring in some certain things that might be useful for example world models
但引入一些可能有用的事情会很好,例如世界模型
there's growing excitement on that do you think there will be any use in this coming year for world models in the llm space
对此有越来越多的兴奋,你认为在未来一年里,世界模型在 LLM 领域会有任何用途吗
yes i do think so
是的,我确实这么认为
also with llms what's an interesting thing here is i think if we unlock more llm capabilities
另外,对于 LLM,这里有趣的一点是,我认为如果我们解锁更多 LLM 能力
it also automatically unlocks all the other fields because or not unlocks but like makes progress faster because you know a
它也会自动解锁所有其他领域,因为,或者说不是解锁,而是像让进展更快,因为你知道一个
Lot of researchers and engineers use LLMs like we said for coding, so
很多研究人员和工程师像我们说的那样用LLM来编程,所以
even if they work on robotics, if you like optimize these LMs that help you with coding, you know, it's like it pays off
即使他们做机器人,如果你优化这些帮你编程的语言模型,你知道,这就像是有回报的
pays off 常用搭配
有回报,值得
表示投入时间或努力后获得好的结果。
but then, uh, yes, the world models are interesting
但然后,呃,是的,世界模型很有趣
it's basically where you have the model run a simulation of the world in a sense, like a little toy thing of the real thing
基本上就是让模型运行一个世界模拟,某种意义上,就像真实事物的一个小玩具
which can again unlock capabilities, um like that are the LMs not aware of, is can simulate things
这又能解锁一些能力,嗯,就像语言模型没有意识到的,就是能模拟事物
and I think, see, this is like something I think LMs, they just happen to work well by pre-training and then doing the next token prediction
我认为,看,这就像我觉得语言模型,它们只是碰巧通过预训练然后做下一个词预测就工作得很好
but we could do this even a bit, you know, like sophisticated in a sense
但我们可以做得更复杂一点,你知道,某种意义上
so what I'm saying is like with there's like a, I think it was by Meta, a paper Coda world models, um, so where they basically apply the concept of world models to LLMs
所以我想说的是,有一个,我觉得是Meta的论文Coda世界模型,嗯,他们基本上把世界模型的概念应用到LLM上
again, where they, and so instead of just having next token prediction and verifiable rewards checking the answer correctness, they also make sure the intermediate variables are correct
再次,他们,所以不只是有下一个词预测和可验证奖励检查答案正确性,他们还确保中间变量是正确的
you know, like it's kind of like a, the model is learning basically a code environment
你知道,就像这有点像,模型基本上在学习一个代码环境
in a sense and I think this makes a lot of sense
从某种意义上说,我觉得这很有道理
makes a lot of sense 常用搭配
很有道理
表示某事物逻辑上合理或容易理解。
it's just like expensive to do but it is like making things more sophisticated
只是做起来很昂贵,但它确实让东西变得更复杂
like modeling um like modeling the whole thing not just the result
比如建模,嗯,比如对整个东西建模,而不仅仅是结果
and uh so it can add more value.
而且呃,所以它能增加更多价值。
I remember when I was a grad student, there is a... so there's a competition called casp i think
我记得我读研究生的时候,有一个……所以有一个叫CASP的比赛,我想
where they do a protein structure prediction like they predict the structure of a protein that is not solved yet
他们做蛋白质结构预测,就像他们预测一个尚未解决的蛋白质的结构
and at that point so in a sense this is actually great
而在那时,所以从某种意义上说这实际上很棒
and i think we need something like that for lms
我觉得我们需要类似的东西来评估语言模型
also where you do the benchmark but no one does so you hand in the results
也是你进行基准测试但没人做,所以你提交结果
but no one knows the solution and then after the fact someone reveals that
但没人知道解决方案,然后事后有人揭示
after the fact 常用搭配
事后
表示在事情发生之后才做某事或才知道。
but alpha fold when it came out it crushed uh you know this benchmark
但AlphaFold问世时,它碾压了呃你知道这个基准测试
i mean there were also multiple iterations.
我的意思是也有多次迭代。
But I remember the first one, I'm not an expert in that subfield,
但我记得第一个,我不是那个子领域的专家,
but the first one explicitly modeled the physical interactions of the, you know, the physics of the molecule,
但第一个明确地模拟了物理相互作用,你知道,分子的物理,
also like the angles, impossible angles.
还有角度,不可能的角度。
And then in the next version, I think they got rid of this.
然后在下一个版本中,我想他们去掉了这个。
got rid of 常用搭配
去掉、移除
口语中表示摆脱或删除某物,比 remove 更随意。
So, and just with brute force scaling it up.
所以,仅仅靠蛮力扩大规模。
brute force 常用搭配
蛮力;靠简单粗暴地堆资源
指不靠精巧方法,而是靠大量算力或重复尝试硬解决问题。
And I think with LLMs, we are currently in this brute force scaling because it just happens to work.
我认为对于大语言模型,我们目前正处于这种蛮力扩展阶段,因为它恰好有效。
brute force scaling 常用搭配
蛮力式扩展
用于描述靠单纯增加算力或数据规模来提升效果,而非改进方法。
But I do think also at some point it might make sense to bring back this thing.
但我确实认为,在某个时候,把这种东西重新引入可能是有意义的。
make sense 常用搭配
有道理、合理
表示某做法在逻辑上说得通,常用于讨论是否值得做某事。
bring back 常用搭配
重新引入、恢复
指把以前有过但后来去掉的东西重新拿回来使用。
And I think with world models, I think that is where I think that might be actually quite cool.
我认为对于世界模型,我觉得那可能实际上相当酷。
I mean, yeah.
我是说,是的。
And of course, also for robotics.
当然,对机器人技术也是如此。
That is completely unrelated from MLMs.
那和MLM完全无关。
Yeah, yeah. And robotics is very explicit.
是的,是的。机器人技术非常明确。
So there's the problem of locomotion or manipulation.
所以存在运动或操作的问题。
Locomotion is much more solved, especially in the learning domain.
运动问题解决得更好,尤其是在学习领域。
But there's a lot of value, just like with the initial protein folding systems, bringing in the traditional model-based methods.
但有很多价值,就像最初的蛋白质折叠系统一样,引入传统的基于模型的方法。
so you don't it's it's unlikely that you can just learn the manipulation or the whole body local manipulation problem end to end that's the dream but then you realize when you look at the magic of the human hand and the complexity of the real world you realize it's really hard to learn this all the way through the way i guess alpha fold two did i'm
所以你不……不太可能直接端到端地学会操作或整个身体局部操作问题,那是梦想,但当你看到人类手部的神奇和现实世界的复杂性时,你会意识到,要像AlphaFold2那样一路学到底真的很难,我
end to end 常用搭配
端到端地
指从输入到输出整个过程一次性完成,中间不需要人工分步处理。
all the way through 常用搭配
自始至终、全程
强调从头到尾完整地完成某过程,常用于否定句表示难以全程做到。
excited about the robotic learning space so i think it's collectively getting like supercharged by all the excitement and
对机器人学习领域感到兴奋,所以我认为它正被所有的兴奋所集体推动,变得像被超级充电一样,而且
Investment in language models generally, where they're getting like the infrastructure for training transformers,
一般来说,对语言模型的投资,也就是他们获得用于训练Transformer的基础设施,
which is like a general modeling thing, is becoming like world-class industrial tooling,
这就像一种通用的建模工具,正在变成世界级的工业级工具,
where anything wherever that was a limitation for robotics, it's just like way better.
凡是以前对机器人技术构成限制的地方,现在都好多了。
There's where more compute and then on top of like they take these language models and use them as kind of central units
有了更多算力,然后在此基础上,他们把这些语言模型当作核心单元来使用,
where you can do interesting explorative work around something that kind of already works,
在这里你可以围绕某种已经能用的东西做有趣的探索性工作,
and then I see it emerging as like kind of like we talked about Hugging Face Transformers and Hugging Face.
然后我看到它正在兴起,就像我们谈到的Hugging Face Transformers和Hugging Face那样。
I think when I was at Hugging Face, I was trying to get this to happen, but it was too early.
我想我在Hugging Face的时候,我试图促成这件事,但当时太早了。
It's like these open robotic models on Hugging Face and be having people be able to contribute data and fine-tune them.
就像在Hugging Face上有这些开放的机器人模型,让人们能够贡献数据并微调它们。
I think we're much closer now that the investment in robotics and, I think, self-driving cars is related and enables this,
我认为我们现在近多了,因为对机器人技术的投资,以及我认为对自动驾驶汽车的投资,是相关的,并且促成了这一点,
where it's like once you get to the point where you can have this sort of ecosystem
就像一旦你到了能够拥有这种生态系统的地步,
where somebody can download a robotics model and maybe fine-tune it to their robot, or share data sets across the world, and there's some data
在这里,有人可以下载一个机器人模型,也许针对他们的机器人进行微调,或者在世界各地共享数据集,而且有一些数据
There's some work in this area, like RTX, I think, is a few years ago
这个领域有一些工作,比如RTX,我想,是几年前
where people are trying to do that, but I think once they have this ecosystem
人们试图做这件事,但我认为一旦他们有了这个生态系统
it'll look very different. And then this whole post-ChatGPT boom is putting more resources into that
它看起来会非常不同。然后整个后ChatGPT热潮正在向其中投入更多资源
which I think is a very good area for doing research. This is also
我认为这是一个非常好的研究领域。这也
is also resulting in much better, more accurate, more realistic simulators being built,
也导致构建出更好、更准确、更逼真的模拟器,
closing the sim-to-real gap in the robotics space.
缩小了机器人领域的模拟到现实的差距。
sim-to-real gap 常用搭配
模拟与现实之间的差距
机器人领域术语,指模拟环境中训练的效果迁移到真实世界时的落差。
But, you know, you mentioned a lot of excitement in the robotics space and a lot of investment.
但是,你知道,你提到了机器人领域的很多兴奋和大量投资。
The downside of that, which happens in hype cycles,
其缺点,在炒作周期中会发生,
I personally believe most robotics people believe that it's not, robotics is not going to be solved at the time scale as being kind of implicit or explicitly promised.
我个人认为大多数机器人领域的人认为,它不会,机器人不会在那种隐含或明确承诺的时间尺度上被解决。
And so what happens when there's all these robotics companies that spring up and then they don't have a product that works,
所以当所有这些机器人公司涌现出来,然后他们没有可行的产品时,会发生什么,
spring up 常用搭配
涌现、迅速出现
形容公司、事物等短时间内大量冒出来。
then there's going to be this kind of crash of excitement, which is nerve-wracking.
然后就会出现这种兴奋的崩溃,这让人紧张。
nerve-wracking 常用搭配
让人紧张不安的
口语中形容某事令人焦虑、提心吊胆。
Hopefully something else will come in and keep swooping in so that the continued development of some of these ideas keeps going.
希望其他东西会介入并不断涌入,以便其中一些想法的持续发展能够继续下去。
swooping in 常用搭配
突然介入、涌入
口语中形容某事物或人迅速进入并接管局面。
I think it's also related to the continual learning issue, essentially,
我认为这也与持续学习的问题有关,本质上,
where the real world is so complex where with LLMs,
现实世界如此复杂,而对于大语言模型,
Yeah, you don't need to really have something learn for the user
是的,你不需要真正为用户学习某些东西,
because there are a lot of things everyone has to do.
因为有很多事情每个人都要做。
Everyone maybe wants to fix their grammar in their email or code or something like that.
每个人可能都想修正他们电子邮件或代码或类似东西中的语法。
It's more constrained.
这更受限制。
So you can kind of prepare the model for that.
所以你可以为那个准备模型。
But preparing the robot for the real world, that's harder.
但为现实世界准备机器人,那更难。
I mean, you have the foundation models, the robotic foundation models,
我的意思是,你有基础模型,机器人基础模型,
but you can learn certain things like grasping things.
但你可以学习某些事情,比如抓取东西。
But then again, I think everyone's house is different.
但话说回来,我认为每个人的房子都不同。
You know, like it's so different.
你知道,就像它是如此不同。
And that is, I think, where the robot would have to learn on the job, essentially.
而那,我认为,就是机器人必须边干边学的地方,本质上。
on the job 常用搭配
在实际工作中、边干边学
指在实际工作过程中学习或进行,而非事先培训。
And I think that, I guess, is the bottleneck right now,
而我认为,我猜,那就是现在的瓶颈,
like how to, you know, customizing it on the fly, essentially.
比如如何,你知道,即兴定制它,本质上。
on the fly 常用搭配
即兴地、临时地
指在做某事的过程中即时调整或处理,不提前准备。
I don't think I can possibly understate the importance of the thing that doesn't get talked about almost at all by robotics folks or anyone is safety.
我认为我无论如何强调安全的重要性都不为过,而机器人领域的人或任何人几乎完全不谈论它。
understate the importance 常用搭配
低估重要性
常用于 I can't understate... 等表达,强调某事极其重要。
All the interesting complexities we talk about learning, all the failure modes and failure cases, everything we've been talking about at LLM,
我们谈论学习时提到的所有有趣的复杂性,所有的失败模式和失败案例,我们在LLM上一直在谈论的一切,
sometimes it fails in interesting ways.
有时它会以有趣的方式失败。
All of that is fun and games in the LLM space.
所有这些在LLM领域都是好玩的事。
fun and games 常用搭配
轻松好玩的事
口语中常暗示某事在别处无伤大雅,但在严肃场合就不行了。
In the robotic space, in people's homes, across millions of minutes, billions of interactions,
在机器人领域,在人们的家中,跨越数百万分钟、数十亿次互动,
you really are almost allowed to fail never.
你真的几乎永远不被允许失败。
When you have embodied systems that are put out there in the real world,
当你有实体系统被放到现实世界中时,
you just have to solve so many problems you never thought you'd have to solve
你只需要解决那么多你从未想过要解决的问题,
when you're just thinking about the general robot learning problem
当你只是在思考通用机器人学习问题时
i'm so bearish on in-home learned robots for consumer purchase
我非常看空用于消费者购买的居家学习型机器人
bearish on 常用搭配
看空、不看好
金融和口语中表示对某事物前景持悲观态度。
i'm very bullish on self-driving cars and i'm very bullish for robotic automation
我非常看好自动驾驶汽车,也非常看好机器人自动化
bullish on 常用搭配
看好、看涨
金融和口语中表示对某事物前景持乐观态度。
eg like amazon distribution where amazon has built whole new distribution centers designed for robots first rather than humans
例如亚马逊的配送,亚马逊建造了全新的配送中心,首先是为机器人而非人类设计的
there's a lot of excitement in ai circles about ai enabling automation and like mass scale manufacturing
在AI圈子里,人们对AI赋能自动化和大规模制造感到非常兴奋
and i do think that the path to robots doing that is more reasonable
我确实认为机器人做到这一点的路径更合理
where it's like a thing that is designed and optimized to do a repetitive task
它就像是一个被设计和优化来做重复性任务的东西
that a human could conceivably do but doesn't want to and then
人类可以想象去做但不想做,然后
i'm but so
我是,但是,所以
but it's also going to
但它也将会
take a lot longer than people probably predict.
花费的时间可能比人们预测的要长得多。
I think that the leap from AI singularity to we can now scale up mass manufacturing in the U.S. because we have a massive AI advantage
我认为,从AI奇点到我们现在能够扩大美国的大规模制造业,因为我们拥有巨大的AI优势,这个跨越
is one that is troubled by a lot of political and other challenging problems.
是一个受到许多政治和其他挑战性问题困扰的跨越。
Let's talk about timelines, uh specifically timelines to AGI or ASI.
我们来谈谈时间线,呃,特别是到AGI或ASI的时间线。
Is it fair, like, as a starting point to say that nobody really agrees on the definitions of AGI and ASI?
作为一个起点,说没有人真正同意AGI和ASI的定义,这公平吗?
I kind of think there's a lot of disagreement, but among, I've been getting pushback where a lot of people kind of say the same thing,
我有点觉得有很多分歧,但是,我最近遇到一些反对意见,很多人说的其实是一回事,
getting pushback 常用搭配
遭到反对、遇到反驳
指提出观点后受到他人的反对或质疑。
which is like a thing that could reproduce most digital economic work, so like the remote worker is a fairly reasonable example.
就是像一种能够复制大多数数字经济工作的东西,所以像远程工作者就是一个相当合理的例子。
And I think OpenAI's definition is somewhat related to that, which is like an AI that can do a lot of economic, like a certain number of economically valuable tasks, which I don't really love as a definition,
我认为OpenAI的定义与此有些相关,就是像一种能够完成很多经济任务,比如一定数量的有经济价值的任务的AI,我不太喜欢这个定义,
somewhat related to 常用搭配
与……有些关联
用于表示两个事物之间存在一定但非完全的联系,语气较委婉。
but I think it could be a grounding point, because language models today,
但我认为它可以作为一个基准点,因为今天的语言模型,
grounding point 常用搭配
基准点;立足点
用于讨论抽象概念时,指可以作为参照或出发点的具体事物。
while immensely powerful, are not this remote worker drop-in, and there are things that you could think of that are could
虽然非常强大,但并不是这种远程工作者的替代品,而且你可以想到一些事情,可能
drop-in 常用搭配
即插即用的替代品
形容某物可以直接替换现有事物而无需额外调整,常用于技术或工作场景。
be done by an AI that are way harder than remote work,
由AI来完成,这些任务比远程工作难得多,
which are like solving a finding an unexpected scientific discovery that you couldn't even pause it,
这就像解决一个发现一个意想不到的科学发现,你甚至无法暂停它,
which would be an example of something that somebody says it's like an artificial super intelligence problem,
这会是某人说它像一个超级人工智能问题的例子,
or like taking in all medical records and finding linkages across certain illnesses that people didn't know,
或者像输入所有医疗记录,发现人们不知道的某些疾病之间的关联,
or figuring out that some common drug can treat some niche cancer like they would say that that is like a super intelligence thing,
或者弄清楚某种常见药物可以治疗某种罕见癌症,就像他们会说那是一种超级智能的事情,
so these are kind of natural tears, my problem with it is that it becomes deeply entwined with like the quest for meaning of AI and this religious aspects to it,
所以这些是某种自然的撕裂,我对它的意见是它变得与AI的意义探索以及其中的宗教方面深深交织在一起,
deeply entwined with 常用搭配
与……深深交织在一起
形容两个事物紧密关联、难以分开,常用于抽象概念或复杂关系。
so there's kind of different, there's different paths you can take it and I don't even know if the remote work is a good definition,
所以有某种不同的,有不同的路径你可以走,我甚至不知道远程工作是否是一个好的定义,
because what exactly is that, it's like perfect tool use, I actually I mean I like I don't know if you like the originally titled AI27 report,
因为那到底是什么,它就像完美的工具使用,我实际上我的意思是我喜欢我不知道你是否喜欢最初标题为AI27报告的那个,
they focus more on code and research taste.
他们更关注代码和研究品味。
So the target there is the superhuman coder.
所以那里的目标是超级人类程序员。
So they have several milestone systems, superhuman coder, superhuman AI researcher, then superintelligent AI researcher, and then the full ASI, artificial superintelligence.
所以他们有几个里程碑系统:超人程序员、超人AI研究员,然后是超级智能AI研究员,最后是完整的ASI,人工超级智能。
But after you develop the superhuman coder, everything else falls quickly.
但一旦你开发出超人程序员,其他一切都会迅速实现。
There, the task is to have a fully autonomous, like automate coding.
在那里,任务是拥有一个完全自主的,比如自动化编程。
So any kind of coding you need to do in order to perform research is fully automated.
所以,为了进行研究而需要做的任何编程都是完全自动化的。
And from there, humans would be doing AI research together with that system,
从那时起,人类将与那个系统一起进行AI研究,
and they would quickly be able to develop a system that actually can do the research for you.
他们很快就能开发出一个真正能为你做研究的系统。
That's the idea.
这就是那个想法。
And then initially their prediction was 2027, 28,
然后最初他们的预测是2027年、28年,
Now they've pushed it back by three to four years to 2031, mean prediction.
现在他们将其推迟了三到四年,到2031年,平均预测。
pushed it back 常用搭配
推迟了它
用于表示将计划、日期或预测延后。
Probably my prediction is even beyond 2031.
可能我的预测甚至超过2031年。
But at least you can in a concrete way think about how difficult it is to fully automate programming.
但至少你可以具体地思考完全自动化编程有多难。
Yeah, I disagree with some of their presumptions and dynamics on how it would play out.
是的,我不同意他们关于事情会如何发展的一些假设和动态。
play out 常用搭配
发展;展开
用于描述事情如何逐步发生或演变,常用于讨论计划、情景或事件。
But I think they did good work in the scenario defining milestones
但我认为他们在定义里程碑的情景中做得很好
that are concrete and to tell a useful story,
这些里程碑是具体的,并且能讲述一个有用的故事,
which is why the reach for this AI 2027 document well transcended
这就是为什么这份AI 2027文件的影响远远超越了
Silicon Valley is because they told a good story
硅谷,因为他们讲了一个好故事
and they did a lot of rigorous work to do this I think
而且他们做了大量严谨的工作来实现这一点,我认为
the camp that I fall into is that like AI is like so-called jagged
我所属的阵营是,AI就像所谓的锯齿状
which will be excellent at some things and really bad at some things
它在某些事情上会非常出色,在某些事情上却非常糟糕
I think that when they're close to this automated software engineer
我认为当它们接近这个自动化软件工程师时
what it will be good at is that traditional ML systems
它擅长的是传统的机器学习系统
and front end the model is excellent at
以及前端,模型非常擅长
but the distributed ML, the models are actually really quite bad at
但分布式机器学习,模型实际上非常不擅长
because there's so little training data on doing large-scale distributed learning and things.
因为关于进行大规模分布式学习等方面的训练数据太少了。
And that's something that we already see.
而这是我们已经看到的现象。
And I think those are just getting amplified.
我认为这些现象正在被放大。
And then it's kind of messier in these trade-offs.
然后在这些权衡中,情况就有点更混乱了。
And then there's like, how do you think AI research works and so on?
然后还有,比如,你认为AI研究是如何运作的等等?
So you think basically superhuman coder is almost unachievable,
所以你认为基本上超级人类程序员几乎无法实现,
meaning like because of the jagged nature of the thing,
意思是,因为这件事的锯齿状特性,
you're just always going to have gaps in capabilities.
你总是会在能力上存在差距。
I think it's assigning completeness to something where the models are kind of superhuman at some types of code.
我认为这是把完整性赋予某些模型在某些代码类型上已经达到超人水平的东西。
And I think that will continue.
而且我认为这种情况会持续下去。
And people are creative, so they'll utilize this incredible abilities to fill in the weaknesses of the models and move really fast.
而人们很有创造力,所以他们会利用这些惊人的能力来填补模型的弱点,并快速推进。
fill in the weaknesses 常用搭配
弥补弱点
用于表示补足某事物的不足之处,常用于讨论能力或系统的缺陷。
There will always kind of be this, I've seen for a long time, this dance between the humans are enabling this thing that the model can't do.
总会存在这种,我长期以来看到的,人类在促成模型做不到的事情之间的这种共舞。
And the best AI researchers are the ones that can enable this superpower.
而最优秀的AI研究人员就是那些能够促成这种超能力的人。
And I think this aligns to what we already see.
我认为这与我们已经看到的情况一致。
I think like cloud code for building a website.
我觉得就像用云代码来搭建网站。
You can stand up a beautiful website in a few hours or do data analysis.
你可以在几个小时内搭建一个漂亮的网站,或者做数据分析。
stand up 常用搭配
搭建;建立
用于表示快速创建或启动某个系统、网站或服务。
And I don't think it's going to keep getting better at these things.
而且我认为它不会在这些事情上持续变得更好。
It'll pick up some new code skills and stuff that it'll get along the way.
它会在此过程中学会一些新的代码技能和东西。
And kind of linking to what's happening in big tech is like this AI 2027 report is like
而联系到大型科技公司正在发生的事情,就像这份AI 2027报告一样
it leans into the singularity idea where I think research is messy and social and largely in the data in ways that AI models can't process.
它倾向于奇点理念,而我认为研究是混乱的、社会性的,并且很大程度上存在于数据中,而AI模型无法处理这些。
leans into 常用搭配
倾向于;偏向
用于表示某事物倾向于某种理念或方向,常含主动拥抱之意。
But like what we do have today is really powerful in these,
但就像我们今天所拥有的,在这些方面确实很强大,
tech companies are all collectively buying into this with tens of billions of dollars of investment
科技公司都在集体投入其中,投入数百亿美元的资金
buying into 常用搭配
认同并投入
用于表示接受并支持某个想法或趋势,常涉及资金或行动上的投入。
so like we are going to get some much better version of chat gpt a much better version of cloud code than we already have
所以就像我们会得到一些更好的ChatGPT版本,比我们已有的更好的Claude Code版本
I think that it's just like hard to predict where that is going
我认为只是很难预测那会走向何方
but the like bright clarity of that future is why some of the most powerful people in the world are putting so much money into this
但那个未来的光明清晰,正是为什么世界上一些最有权势的人正在向这个领域投入如此多的资金
and I think it's just kind of small differences between like we don't actually know what a better version of chat gpt is
而且我认为这只是些微小的差异,就像我们实际上并不知道更好的ChatGPT版本是什么
but also like can it automate AI research I would say probably not at least in this time frame
但也比如它能否自动化AI研究,我会说可能不会,至少在这个时间范围内
like big tech is going to spend 100 billion dollars much faster than we get a automated AI researcher
就像大型科技公司会花费1000亿美元,比我们获得自动化AI研究员要快得多
that enables a AI research singularity so you think your prediction would be what like if this is even a useful milestone
那能实现AI研究奇点,所以你认为你的预测会是什么,如果这甚至是一个有用的里程碑
we're more than 10 years out I would say less than that on the software side,
我们还有超过10年的时间,我会说在软件方面不到那么久,
but I think longer than that on the things like research.
但我认为在像研究这样的事情上会更久。
It's just like, for fun, try to imagine a world where all software writing is fully automated.
就像,为了好玩,试着想象一个所有软件编写都完全自动化的世界。
Can you imagine that world?
你能想象那个世界吗?
By the end of this year, the amount of software that'll be automated will be so high.
到今年年底,将被自动化的软件数量会非常高。
But it'll be the things of you're trying to train a model with RL and you need to have multiple bunches of GPUs communicating with each other, that'll still be hard, but I think it'll be much easier.
但那些你试图用强化学习训练模型、需要多组GPU相互通信的事情,仍然会很难,但我认为会容易得多。
One of the ways to think about this, the full automation of programming, is just think of lines of useful code written, the fraction of that to the number of humans in the loop.
思考这个问题的一种方式,即编程的完全自动化,就是想想所编写的有用代码行数,与循环中人类数量的比例。
So presumably, there'll be for a long time humans in the loop of software writing, it'll just be fewer and fewer relative to the amount of code written, right?
所以大概在很长一段时间内,软件编写循环中仍会有人类,只是相对于所编写的代码量会越来越少,对吧?
And the SC superhuman coder, I think the presumption there is it goes to zero, the number of humans in the loop.
而SC超级人类编码员,我认为那里的假设是循环中人类的数量会降到零。
What does that world look like when the number of humans in the loop is in the hundreds, not in the hundreds of thousands?
当循环中人类的数量是几百,而不是几十万时,那个世界会是什么样子?
I think software engineering will be driven more to system design and goals of outcomes, where I do think software is largely going to be.
我认为软件工程将更多地转向系统设计和结果目标,我确实认为软件将主要朝着这个方向发展。
I think this has been happening over the last few weeks
我认为这在过去几周里一直在发生
where people have gone from a month ago of like, oh, yeah, agents are kind of slop,
人们从一个月前的“哦,是的,代理有点垃圾”
which is a famous Carpathie quote to like the what is a little bit of a meme of like the industrialization of software
这是卡帕西的一句名言,有点像软件工业化的梗
when anyone can just create software at their fingerprints.
当任何人都能仅凭指纹创造软件时。
Like, I do think we are closer to that side of things
就像,我确实认为我们更接近那一面了
closer to that side of things 常用搭配
更接近那种情况或方向
表示某个趋势或状态正在向某个方向靠近,口语中常用。
and it takes direction and like understanding how the systems work to extract that best from the language models.
并且需要指导和理解系统如何工作,才能从语言模型中提取出最好的结果。
extract that best from 常用搭配
从……中提取出最好的部分
用于描述从某事物中获取最大价值或最佳结果。
And I think it's hard to like accept the gravity of how much is going to change with software development
而且我认为很难接受软件开发将会发生多大变化的严重性
accept the gravity of 常用搭配
接受……的严重性
用于强调认识到某事的重大或严肃程度。
and how many more people can do things without ever looking at it.
以及有多少更多人可以在完全不看它的情况下做事。
without ever looking at it 常用搭配
完全不用看它
强调不需要直接查看或接触某物就能完成事情。
I think what's interesting is to think about whether these systems will be independent,
我认为有趣的是思考这些系统是否会独立,
like completely independent in the sense that, well, I have no doubt that LLMS will kind of at some point solve coding in a sense like calculators.
就像完全独立,意思是,嗯,我毫不怀疑LLM会在某个时候像计算器一样解决编码问题。
I have no doubt that 句型
我毫不怀疑……
I have no doubt that [clause]
用于表达对某事非常确信。
solve calculating, right? So at some point humans developed a tool that, you know, you never need a human to calculate that number, you just type it in and it's an algorithm.
解决计算问题,对吧?所以到了某个时候,人类发明了一种工具,你知道,你不再需要人来计算那个数字,你只要输入进去,它就是一个算法。
at some point 常用搭配
在某个时候
指不确定的具体时间点,常用于叙述未来或过去的事件。
You, you can do it in a, in that sense, and I I think that's the same probably for coding.
你,你可以在那个意义上做到这一点,而且我,我觉得编程大概也是一样的。
in that sense 常用搭配
在那个意义上
用于限定或澄清前面所说的意思。
But the question is, so I think what will happen is, yeah, you will just say build that website, it will make a really good website, and then you maybe refine it.
但问题是,所以我觉得会发生的是,是的,你只要说建那个网站,它就会做出一个非常好的网站,然后你也许再完善它。
what will happen is 句型
将会发生的是……
what will happen is [clause]
用于引出对未来的预测或解释。
But will it do things independently, where, so will you be still having humans asking the AI to do something?
但它会独立地做事吗,也就是说,你还会让人来要求AI做某件事吗?
Like, will there be a person, say, build that website? Or will there be AI that just builds websites or something or whatever?
比如,会不会有一个人说,建那个网站?还是会有AI直接就能建网站之类的,或者随便什么?
I think using, talking about building websites is the... Too simple.
我觉得用,谈论建网站是……太简单了。
The problem with websites and the problem with the web, you know, HTML and all that kind of stuff, it's very resilient to just slop.
网站的问题,还有网络的问题,你知道,HTML和所有那类东西,它对随便糊弄的东西非常宽容。
all that kind of stuff 常用搭配
诸如此类的东西
口语中用于列举类似事物,表示不一一详述。
It will show you slop, as good as showing slop.
它会给你展示糊弄的东西,展示糊弄的东西也一样好。
I would rather think of safety critical systems like asking AI to end-to-end generate something that manages logistics or manages cars, a fleet of cars, all that kind of stuff.
我更愿意想想安全关键系统,比如让AI端到端地生成某种管理物流的东西,或者管理汽车、一个车队,所有那类东西。
I would rather think of 句型
我更愿意考虑……
I would rather think of [something]
用于表达偏好,提出另一种思考角度。
So end-to-end generates that for you. I think a more intermediate example is take something like Slack or Microsoft Word.
所以端到端地为你生成那个。我觉得一个更中间的例子的就是拿像Slack或Microsoft Word这样的东西。
I think if the organizations allow it, AI could very easily implement features end-to-end and do a fairly good job for things that you want to try.
我认为如果各个组织允许的话,AI可以非常轻松地端到端实现功能,并且对于你想尝试的事情做得相当不错。
do a fairly good job 常用搭配
做得相当不错
用于评价某人在某事上表现良好。
you want to add a new tab in Slack that you want to use.
你想在Slack里添加一个你想用的新标签页。
And I think AI will be able to do that pretty well.
而且我认为AI能够做得相当好。
Actually, that's a really great example.
实际上,那真是个很好的例子。
How far away are we from that? Like this year?
我们离那还有多远?比如今年?
See, I don't know. I don't know.
你看,我不知道。我不知道。
I guess I don't know how bad production code bases are.
我猜我不知道生产代码库有多糟糕。
But I think that within like on the order of low years, a lot of people are going to be pushed to be more of like a designer and product manager
但我认为,在大约几年之内,很多人会被推向更像设计师和产品经理的角色
on the order of 常用搭配
大约,在……数量级
用于表示大概的数量或范围,较正式。
where you have multiple of these agents that can try things for you.
在那里你有多个这样的代理可以为你尝试各种事情。
And they might take one to two days to implement a feature or attempt to fix a bug.
它们可能需要一到两天来实现一个功能或尝试修复一个bug。
And you have these dashboards, which I think Slack is actually a good dashboard where your agents will talk to you.
而且你有这些仪表盘,我认为Slack实际上是一个很好的仪表盘,你的代理会在那里和你交流。
And you'll then give feedback.
然后你会给出反馈。
But things like, I make a website, it's like, do you want to make a logo that's passable?
但像这样的事情,我做一个网站,它就像,你想做一个还过得去的标志吗?
Like, I think these like cohesive design things and this style is going to be very hard for models and deciding on what to add at the next time.
就像,我觉得这些连贯的设计元素和这种风格,对模型来说会非常难,还有决定下次要加什么。
I just, okay, so I hang out with a lot of programmers, and some of them are a little bit on the skeptical side in general.
我只是,好吧,所以我和很多程序员混在一起,他们中有些人总体上有点偏怀疑态度。
hang out with 常用搭配
和……一起玩,相处
口语中表示与某人共度休闲时间。
That's just vibe-wise, they're like that.
那只是感觉上,他们就是那样。
I just think there's a lot of complexity involved in adding features to complex systems.
我只是觉得,给复杂系统添加功能涉及很多复杂性。
Like if you look at the browser, Chrome, if I wanted to add a feature,
就像如果你看看浏览器,Chrome,如果我想添加一个功能,
if I wanted to have tabs as opposed to up top, I want them on the left side.
如果我想让标签页不在顶部,而是放在左侧。
as opposed to 常用搭配
而不是,与……相对
用于对比两个不同的事物或选择。
Interface, right?
界面,对吧?
I think we're not, it's not a next year thing.
我觉得我们不是,这不是明年的事。
One of the Cloud releases this year, one of their tests was we give it a piece of software and leave Cloud to run to recreate it entirely.
今年Cloud发布的一个版本,他们的测试之一是我们给它一个软件,然后让Cloud运行来完全重建它。
And it can already almost rebuild like Slack from scratch, just given the parameters of the software and left in a sandbox environment.
而且它已经几乎可以从零开始重建像Slack这样的东西,只要给出软件的参数,然后把它放在沙盒环境里。
from scratch 常用搭配
从零开始
表示从头做起,不借助已有的东西。
So from the scratch part, I like almost better.
所以从零开始这部分,我几乎更喜欢。
So it might be that the smaller, newer companies are advantaged and they're like, we don't have to have the bloat and complexity.
所以可能是这样,规模更小、更新的公司有优势,他们会说,我们不必有臃肿和复杂性。
and therefore this future exists.
因此这个未来是存在的。
And I think this gets to the point that you mentioned
我认为这就涉及到了你提到的那个点
gets to the point 常用搭配
涉及到要点
用于引出核心问题或关键点。
that some people you talk to are skeptical.
即你交谈过的一些人是持怀疑态度的。
And I think that's not because the LLM can't do X, Y, Z.
我认为这并不是因为LLM不能做X、Y、Z。
It's because people don't want it to do it this way.
而是因为人们不希望它这样做。
Some of that could be a skill issue on the human side.
其中一些可能是人类方面的技能问题。
Unfortunately, we have to be honest with ourselves.
不幸的是,我们必须对自己诚实。
be honest with ourselves 常用搭配
对自己诚实
用于承认事实,不逃避真相。
And some of that could be an under-specification issue.
其中一些可能是规格不足的问题。
So programming, like you're like, you're just assuming, this is like in relationships, and friendships, communication type of issue.
所以编程,就像你只是在假设,这就像在恋爱、友谊、沟通中的那种问题。
You're assuming the LMS somehow is supposed to read your mind.
你在假设LMS不知怎么就应该能读懂你的心思。
read your mind 常用搭配
读懂你的心思
表示猜测或理解别人的想法,常用于否定或假设。
I think this is where spec-driven design is really important.
我认为这就是规格驱动设计真正重要的地方。
Like, you're just using natural language, specify, like, what you want.
就像,你只是用自然语言,指定,比如,你想要什么。
I think that's like, if you talk to people at the labs, they use these in their training and production code.
我认为这就像,如果你和实验室的人交谈,他们在训练和生产代码中使用这些。
Like, Cloud Code is built with Cloud Code, and they all use these things extensively,
就像,Cloud Code是用Cloud Code构建的,他们都广泛使用这些东西,
and Dario talks about how much of Cloud's code owns.
而Dario谈到Cloud的代码有多少是它自己拥有的。
And it's like, these people are slightly ahead in terms of the capabilities they have,
就像,这些人在他们拥有的能力方面稍微领先一些,
in terms of 常用搭配
在……方面
用于引出讨论的具体方面或角度。
they probably spend on inference they could spend 10 to 100 plus x as much as we're spending like
他们可能在推理上花费,他们可能花费10到100多倍于我们花费的,
we're on a lowly 100 or 200 dollar a month plan like they truly let it rip and I think
我们只是在一个区区100或200美元每月的计划上,就像他们真正放手一搏,我认为
let it rip 地道口语
放手一搏,尽情发挥
口语中表示不加限制地全力去做某事。
that that like with the pace of progress that we have it seems like like
那个,就像以我们拥有的进步速度来看,似乎就像
where a year ago we didn't have cloud code and we didn't really have reasoning models and it's like
一年前我们还没有云代码,我们还没有真正的推理模型,就像
the difference between sitting here today and what we can do with these models and it seems like
今天坐在这里和我们可以用这些模型做的事情之间的区别,似乎就像
there's a lot of like there's a lot of low hanging fruit to improve them
有很多,就像有很多容易改进它们的地方
low hanging fruit 常用搭配
容易实现的目标
比喻容易获得成果或改进的地方。
the failure modes are pretty dumb it's like
失败模式相当愚蠢,就像
Claude you tried to use the CLI command I don't have installed 14 times and then I sent you the command to run
Claude,你尝试使用我没有安装的CLI命令14次,然后我发送给你要运行的命令
it's like that thing from a modeling perspective is pretty fixable
就像从建模的角度来看,那件事是相当可修复的
so I I agree with you I've been uh becoming more and more bullish in general speaking to what you're articulating
所以我同意你的观点,我一直在呃变得越来越看好,总的来说,对于你所阐述的
I think it is a human skill issue so Anthropic is leading the way in or other companies in understanding how to best
我认为这是一个人类技能问题,所以Anthropic在或其他公司在理解如何最好地方面处于领先地位
leading the way 常用搭配
引领潮流,带头
表示在某方面处于领先地位或起带头作用。
use the models for programming, therefore they're effectively using them.
使用这些模型进行编程,因此他们实际上是在使用它们。
I think there's a lot of programmers on the outskirts.
我认为有很多边缘的程序员。
They're like, they don't... I mean, there's not a really good guide on how to use them.
他们就像,他们不……我的意思是,没有真正好的使用指南。
People are trying to figure it out exactly.
人们正在试图弄清楚它到底是怎么回事。
figure it out 常用搭配
弄清楚,弄明白
表示通过思考或尝试找到解决办法。
It might be very expensive, like it might be that the entry point for that is two thousand dollars a month,
它可能非常昂贵,比如入门价格可能是每月两千美元,
which is only tech companies and rich people.
这只有科技公司和富人才负担得起。
Which is like, that could be it.
这就像,可能就是那样。
But it might be worth it.
但这可能值得。
worth it 常用搭配
值得
表示某事物有价值或值得付出。
I mean, if the final result is a working software system, it might be worth it.
我的意思是,如果最终结果是一个可运行的软件系统,那可能就值得。
But by the way, it's funny how we converge from the discussion of timeline to AGI to something more pragmatic and useful.
但顺便说一句,有趣的是我们如何从讨论AGI时间线收敛到更务实和有用的东西。
Is there anything concrete and interesting and useful and profound to be said about timeline to AGI and ASI?
关于AGI和ASI的时间线,有什么具体、有趣、有用且深刻的说法吗?
Or are these discussions a bit too detached from the day-to-day?
还是这些讨论有点太脱离日常生活了?
detached from 常用搭配
脱离,与……无关
表示与某事物没有联系或距离感。
There's interesting bets.
有一些有趣的赌注。
So there's a lot of people trying to do reinforcement learning with verifiable rewards,
所以有很多人尝试用可验证的奖励进行强化学习,
but in real scientific domains where there's startups that are spending, like they have hundreds of millions of dollars of funding and they have wet labs
但在真正的科学领域,有初创公司投入,比如他们有数亿美元的资助,还有湿实验室,
where they're having language models propose hypotheses that are tested in the real world.
在那里他们让语言模型提出假设,并在现实世界中进行测试。
And I would say that I think they're very early or they're early,
我会说,我认为它们还非常早期,或者说它们处于早期,
but with the pace of progress, it's like maybe they're early by six months
但以进步的速度来看,也许它们只是早了六个月,
and they make it because they were there first
而它们能成功是因为它们是最早到那里的,
or maybe they're early by eight years and you don't really know.
或者也许它们早了八年,而你并不真的知道。
So I think that that type of moonshot to branch this momentum into other other sciences is like,
所以我认为,那种把这种势头分叉到其他科学的登月式尝试,就像,
OK, like that would be very transformative if like alpha fold moments happen in all sorts of other scientific domains
好吧,如果类似AlphaFold的时刻在其他各种科学领域发生,那将是非常具有变革性的,
by like a startup solving this.
由像一家初创公司来解决这件事。
I think there are startups, I think maybe Harmonic is one where they're going all in on language models plus lean for math.
我认为有一些初创公司,我想也许Harmonic就是其中之一,他们全力投入语言模型加上用于数学的Lean。
going all in on 常用搭配
全力投入某事
表示把全部资源或精力押注在某件事上,常用于商业或投资语境。
I think you had another podcast guest who talked about this recently and it's like we don't know exactly what's going to fall out.
我记得你最近有另一位播客嘉宾谈到过这个,就像我们并不确切知道会有什么结果。
of spending 100 million dollars on that model and most of them will fail
在那个模型上花费1亿美元,而其中大多数都会失败
but a couple of them might be breakthroughs that are very different than chat gpt or cloud code type software experiences
但其中少数可能会是突破,与ChatGPT或云代码类型的软件体验非常不同
like a tool that's only good for a phd mathematician but makes them 100x effective
就像一个只对博士数学家有用的工具,但能让他们效率提高100倍
like okay well i agree i think this will happen in a lot of domains especially also like domains that have a lot of you know resources like finance and legal and pharmaceutical companies
就像,好吧,我同意,我认为这会在很多领域发生,尤其是那些拥有大量资源的领域,比如金融、法律和制药公司
but then again is it really agi again because we are now specializing it again and then
但话又说回来,这真的是AGI吗?因为我们又在专门化它,然后
but then again 地道口语
不过话又说回来
口语中用来补充一个相反或质疑的观点,缓和前面说过的话。
again is it really that much different from back in the day how we had specialized algorithms i think it's just the same thing more way more sophisticated
再说,这真的和过去我们拥有专门算法的时候有很大不同吗?我认为这只是同样的事情,但方式更加复杂
but i don't know is there a threshold when we call it agi i guess
但我不知道,当我们称之为AGI时,是否有一个阈值,我猜
i think the the real cool thing is here that we have like the foundation models that we can specialize
我认为真正酷的事情是,我们拥有可以专门化的基础模型
i think that that's like the breakthrough at some point right now i think we are not there yet
我认为这就是某个时刻的突破,现在我认为我们还没有达到
because well first it's too expensive but also you know like chat gpt doesn't just give away that chat gpt to customize it i
因为,首先它太贵了,但还有你知道,像ChatGPT不会直接免费提供那个ChatGPT来定制它,我
think once that's going to be true in some way
想想一旦这在某种程度上成为现实
and I think I can imagine this as a business model
我觉得我可以想象这是一种商业模式
that uh ChatGPT OpenAI says at some point
呃 ChatGPT OpenAI 在某个时候说
like hey you know Bank of America 400 million we will do your custom model or something like that
比如嘿你知道美国银行 4 亿我们会为你定制模型或类似的东西
And I think that will be the huge economic value add.
我认为那将是巨大的经济价值增值。
The other thing, though, is also companies.
不过另一件事也是公司。
I mean, right now, what is the differentiating factor?
我的意思是,现在,差异化因素是什么?
I mean, if everyone uses the same LLM, if everyone uses ChatGPD, they will all do the same thing again.
我的意思是,如果每个人都使用相同的 LLM,如果每个人都使用 ChatGPD,他们都会再次做同样的事情。
I mean, then, well, everyone is moving in lockstep, but usually companies,
我的意思是,那么,嗯,每个人都在同步前进,但通常公司,
they want to have a competitive advantage.
他们想要拥有竞争优势。
competitive advantage 常用搭配
竞争优势
商业语境中表示比对手更有利的条件或能力。
And I think there's no way around using some of their private data and experimenting and maybe specializing.
我认为没有办法绕过使用他们的一些私有数据并进行实验,也许专门化。
there's no way around 句型
无法回避、必须面对
there's no way around [doing something]
表示某件事不可避免,只能去做或接受。
It's going to be interesting.
这将会很有趣。
sitting in the pace of progress, it does just feel like things are coming.
坐在进步的速度中,确实感觉事情正在到来。
I don't think the AGI and ASI thresholds are particularly useful.
我不认为 AGI 和 ASI 阈值特别有用。
I think, I guess, the real question, and this takes us to the remote worker thing,
我想,我猜,真正的问题是,这把我们带到了远程工作者的事情上,
is when are we going to see a big obvious leap in economic impact?
是我们什么时候会看到经济影响出现明显的大飞跃?
Because currently there's not been an obvious leap in economic impact of LL models, for example.
因为目前例如 LL 模型的经济影响还没有出现明显的飞跃。
And that's, you know, aside from AGI or ASI or all that kind of stuff,
而且,你知道,撇开AGI或ASI或那一类东西不谈,
aside from 常用搭配
撇开……不谈;除……之外
用来把某个话题暂时排除在外,转而讨论别的重点。
there's a real question of like, when are we going to see a GDP like jump?
有一个很现实的问题,就是,我们什么时候会看到GDP那样的跃升?
Yeah, it's like, what is the GDP made up of?
是啊,就像,GDP是由什么构成的?
Like a lot of it is like financial services.
比如其中很大一部分是像金融服务。
So like, I don't know what this is.
所以就像,我不知道这是什么。
It's just hard for me to think about the GDP bump.
我只是很难去想象GDP的增长。
But like, I'd say that software development becomes valuable in a different way
但就像,我会说软件开发会以一种不同的方式变得有价值,
when you no longer have to look at the code anymore.
当你不再需要去看代码的时候。
So when it is like cloud will make you a small business,
所以当它像是云会帮你做一个小生意,
which is essentially cloud can set up your website, your bank account, your email, and your whatever else.
本质上就是云可以帮你搭建网站、银行账户、电子邮件,还有别的什么。
And you just have to express what you're trying to put into the world.
而你只需要表达你想把什么放进这个世界。
That's not just an enterprise market, but it is a hard...
这不仅仅是一个企业市场,但它是一个很难的……
I don't know how you get people to try doing that.
我不知道你怎么让人们去尝试做这件事。
I guess if ChatGPT can do it, people are trying ChatGPT.
我猜如果ChatGPT能做到,人们就会去尝试ChatGPT。
I think it boils down to the scientific question of how hard is tool use to solve.
我认为这归结为一个科学问题:工具使用有多难解决。
boils down to 常用搭配
归结为
表示复杂问题最终的核心或本质是某一点。
Because a lot of the stuff you're applying, the remote work stuff is tool use.
因为你应用的很多东西,远程工作那些东西,都是工具使用。
It's like how computer use, like how you have an LLM that goes out there,
这就像电脑使用,就像你有一个LLM到外面去,
this agentic system, and does something in the world and only screws up 1% of the time.
这个智能体系统,在现实世界里做事情,只有1%的时候会搞砸。
screws up 常用搭配
搞砸;出错
口语中表示把事情弄糟或犯错,较非正式。
Computer use is a good example of what labs care about
电脑使用是实验室关心的一个很好的例子,
and we haven't seen a lot of progress on.
而我们在这方面还没有看到很多进展。
We saw multiple demos in 2025 of like,
我们在2025年看到了多个演示,比如,
Claude can use your computer or OpenAI had Kua and they all suck.
Claude可以使用你的电脑,或者OpenAI有Kua,但它们都很烂。
So they're also investing money in this.
所以他们也在往这上面投钱。
And I think that will be a good example.
我觉得那会是一个很好的例子。
Whereas that's actually something where it just seems pretty...
而实际上那是某种看起来相当……
Taking over the whole screen seems a lot harder than having an API that they can call in the back end.
接管整个屏幕似乎比有一个他们可以在后端调用的API要难得多。
And some of that is you have to then set up a different environment for the model to work in.
其中一部分原因是,你得为模型设置一个不同的工作环境。
They're not working on your MacBook.
它们不是在你的MacBook上工作。
They are individually interfacing with Google and Amazon and Slack.
它们分别与Google、Amazon和Slack对接。
And they handle all these things in a very different way than humans do.
它们处理所有这些事情的方式和人类非常不同。
So some of those might be structural blockers.
所以其中一些可能是结构性的障碍。
Also, like specification-wise, I think the problem is also for, you know, arbitrary tasks.
另外,从规范的角度来说,我觉得问题也在于,你知道,任意任务。
Well, you still have to specify what you want your LLM to do and how do you do that in a...
嗯,你还是得指定你想让你的LLM做什么,而你要怎么在一个……
What is the environment?
环境是什么?
How do you specify?
你怎么指定?
You can say what the end goal is,
你可以说出最终目标是什么,
but if it can't solve the end goal with LLMs,
但如果它无法用LLM解决最终目标,
if you ask it for text, you can always, you know, clarify, do sub-steps.
如果你让它生成文本,你总是可以,你知道,澄清,做子步骤。
How do you put that information into a system that, let's say, books a travel trip for you?
你如何把这些信息放入一个系统,比如说,为你预订旅行?
you can say, well, you screwed up my credit card information.
你可以说,嗯,你搞砸了我的信用卡信息。
But even to get it to that point,
但即使要达到那个程度,
how do you, as a user, guide the model before it can even attempt that?
作为一个用户,你如何在模型甚至尝试之前引导它?
I think the interface is really hard.
我认为界面真的很难。
Yeah, it has to learn a lot about you specifically.
是的,它必须专门了解你很多。
And this goes to continue learning about the general mistakes that are made throughout and mistakes that are made through you.
这就涉及到继续学习整个过程中犯的一般错误以及通过你犯的错误。
All the AI interfaces are getting set up to ask humans for input.
所有AI界面都在被设置为向人类寻求输入。
I think Cloud Code, we talk about a lot.
我认为Cloud Code,我们经常谈论。
It asks feedback on questions.
它会在问题上征求反馈。
If it doesn't have enough specification on your plan or your desired, it starts to ask questions.
如果它对你的计划或期望没有足够的说明,它就开始提问。
Would you rather?
你更愿意吗?
We talked about memory, which saves across chats, which it's first implementation is kind of odd.
我们谈到了记忆,它跨聊天保存,它的第一个实现有点奇怪。
It'll be like, it'll mention my dog's name or something in a chat.
它会像,它会在聊天中提到我狗的名字之类的。
I'm like, you didn't need to be subtle about this.
我想,你不需要对此这么含蓄。
I don't care.
我不在乎。
But the things that are emerging, our ChatGPT has the pulse feature,
但正在出现的东西,我们的ChatGPT有脉冲功能,
which is like a curated couple paragraphs with links to something to look at or to talk about.
这就像精心挑选的几段话,带有链接,可以看或谈论。
And people talk about how the language models are going to ask you questions,
人们谈论语言模型将如何问你问题,
which I think is a very, it's probably going to work.
我认为这非常,它可能会奏效。
The language model is like, it knows you had a doctor appointment or something.
语言模型就像,它知道你有医生预约之类的。
It's like, hey, how are you feeling after that?
它就像,嘿,之后你感觉怎么样?
Which is like, again, goes into the territory of humans are very susceptible to this
这就像,再次,进入人类非常容易受此影响的领域,
and there's a lot of social change to come.
而且会有很多社会变革到来。
But also like they're experimenting with having the models engaged.
但也像他们在尝试让模型参与进来。
Some people really like this pulse feature,
有些人真的很喜欢这个脉冲功能,
which is it processes your chats and automatically searches for information and puts it in the ChatGPT app.
它会处理你的聊天记录,自动搜索信息并放入ChatGPT应用中。
So there's a lot of things coming.
所以有很多东西即将到来。
I use that feature before and I always feel bad because it does that every day and I rarely check it out.
我以前用过那个功能,我总是感觉不好,因为它每天都那样做,而我很少查看。
It's like how much money, I mean, computers burned on something I don't even look at, you know, where it's like, it's kind of like.
就像多少钱,我是说,电脑烧在我不甚至看的东西上,你知道,就像,有点像。
There's also a lot of idle compute in the world, so don't feel too bad.
世界上也有很多闲置的计算资源,所以别太难过。
Okay. Do you think new ideas might be needed?
好的。你认为可能需要新想法吗?
Is it possible that the path to AGI, whatever that is, however we define that, to solve computer
通往AGI的路径,无论那是什么,无论我们如何定义它,要解决计算机
use more generally, to solve biology and chemistry and physics, sort of the Dario definition of AGI or powerful AI,
更普遍地使用,要解决生物学、化学和物理学,有点像Dario对AGI或强大AI的定义,
do you think it's possible that totally new ideas are needed?
你认为完全新的想法可能是必要的吗?
non-LLM, non-RL ideas, what might they look like?
非LLM、非RL的想法,它们可能是什么样的?
We're not going into philosophy land a little bit.
我们不会稍微进入哲学领域。
For something like a singularity to happen, I would say yes.
对于像奇点这样的事情发生,我会说是的。
And the new ideas can be architectures or training algorithms, which is like fundamental deep learning things.
而新的想法可以是架构或训练算法,这就像是基础深度学习的东西。
But they're in that nature pretty hard to predict.
但它们本质上很难预测。
But I think we won't get very far even without those advances.
但我认为即使没有那些进展,我们也不会走得很远。
like we might get this software solution but it might stop at software and not do computer use without more innovation
就像我们可能得到这个软件解决方案,但它可能止步于软件,没有更多创新就无法进行计算机使用
so i think that it's like a lot of progress will be coming but in if you're going to zoom out like there's still ideas in the next 30 years that are going to look like that was a major like scientific innovation that enabled the next chapter of this and i don't know if it comes in one year or in 15 years yeah i wonder if the bitter lesson holds true for the next hundred years
所以我认为会有很多进展,但如果你要拉远视角,未来30年仍然会有一些想法,看起来像是重大的科学创新,开启了下一章,我不知道它是在一年内还是15年内出现,是的,我想知道苦涩的教训是否在未来一百年仍然适用
zoom out 常用搭配
拉远视角、从宏观层面看问题
讨论复杂话题时,用来表示暂时放下细节、从更大格局来看。
years, what that looks like if scaling laws are fundamental and deep learning
年,如果缩放定律是根本性的,而深度学习
I think the better lesson will always apply, which is compute will become more abundant,
我认为更好的教训将始终适用,那就是计算将变得更加丰富,
but even within abundant compute, the ones that have a steeper scaling law slope or a better offset like this
但即使在丰富的计算中,那些具有更陡的缩放定律斜率或更好的偏移的,就像这样
is a 2D plot of performance and compute, and like even if there's more compute available, the ones that get 100x out of it will win.
是一个性能和计算的二维图,而且即使有更多的计算可用,那些能从中获得100倍收益的将会胜出。
It might be something like literally computer clusters orbiting Earth with solar panels.
它可能就像字面意义上的带有太阳能板的计算机集群环绕地球运行。
The problem with that is heat dissipation.
问题在于散热。
So you get all the radiation from the sun and you don't have any air to dissipate heat.
所以你会接收到来自太阳的所有辐射,而且没有任何空气来散热。
But there is a lot of space to put clusters.
但有很多空间可以放置集群。
There's a lot of solar energy there and you could figure out the heat dissipation.
那里有很多太阳能,你可以解决散热问题。
But there is a lot of energy and there probably could be engineering will to solve the heat problem.
但有很多能源,而且很可能有工程意愿来解决散热问题。
So there could be. Is it possible?
所以有可能。这可行吗?
And we should say that it definitely is possible.
我们应该说这绝对是可能的。
How like that is, uh, is the question that we're basically going to be plateauing this year,
那有多像,呃,是我们今年基本上会趋于平稳的问题,
not in terms of the system capabilities, but what the system capabilities actually mean for human civilization.
不是就系统能力而言,而是系统能力对人类文明实际意味着什么。
So on the coding front, really nice websites will be built.
所以在编程方面,真的会做出很棒的网站。
Very nice autocomplete.
非常好的自动补全。
Very nice way to understand code bases and maybe help debug,
非常好的方式来理解代码库,也许还能帮忙调试,
but really just a very nice helper on the coding front.
但真的只是编程方面一个非常好的帮手。
It can help research mathematicians do some math.
它能帮助研究数学家做一些数学。
It can help you with shopping.
它能帮你购物。
It can help you with, it can help.
它能帮你,它能帮忙。
It's a nice helper.
它是个不错的帮手。
It's clippy on steroids.
它是打了鸡血的Clippy。
What else?
还有什么?
It may be a good education tool and all that kind of stuff,
它可能是个不错的教育工具之类的,
But computer use turns out extremely difficult to solve.
但计算机使用被证明极难解决。
So I'm trying to frame the cynical case in all these domains where there's not a really huge economic impact.
所以我想在这些没有巨大经济影响的领域里,提出那个愤世嫉俗的论点。
We realize how costly it is to train these systems at every level,
我们意识到在每一个层面上训练这些系统有多昂贵,
both the pre-training and the inference,
无论是预训练还是推理,
how costly the inference is, the reasoning, all of that.
推理有多昂贵,推理过程,所有这一切。
Is that possible and how likely is that, do you think?
那可能吗?你觉得可能性有多大?
When you look at the models, there's so much obvious things to improve,
当你看着这些模型时,有太多明显可以改进的地方,
and it takes a long time to train these models and to do this art.
而训练这些模型、做这门艺术需要很长时间。
And it'll take us, with the ideas that we have, multiple years to actually saturate in terms of whatever benchmark or performance we are searching for.
而凭借我们现有的想法,我们实际上需要好几年才能在我们要寻找的任何基准或性能上达到饱和。
It might serve very narrow niches.
它可能只服务于非常狭窄的细分领域。
Like the average ChatGPT 800 million user might not get a lot of benefit out of this,
就像普通的ChatGPT 8亿用户可能不会从中获得很多好处,
but it is going to serve different populations by getting better at different things.
但它会通过在不同方面变得更好来服务不同的人群。
Well, I think what everybody's chasing now is a general system that's useful to everybody.
嗯,我认为现在大家都在追求的是一个对所有人都通用的系统。
So, okay. So if that's not, that can plateau, right?
所以,好吧。所以如果那不是,那可能会停滞,对吧?
I think that dream is actually kind of dying.
我认为那个梦想实际上正在消亡。
kind of dying 常用搭配
有点在消亡、逐渐消失
口语中 kind of 用来弱化语气,表示“有点、某种程度上”,常与形容词或动词连用。
as you talked about with the specialized models where it's like
正如你谈到的专业模型,就像
and multimodal is often like video generation is a totally different thing
而多模态通常就像视频生成是完全不同的东西
that dream is kind of dying is a big statement
那个梦想正在消亡是一个重大的声明
because i don't know if it's dying i don't know if every i don't if you ask the actual financial lab people
因为我不知道它是否在消亡,我不知道是否每个,如果你问实际的金融实验室人员
they i mean they're still chasing it right i do think they are still like rushing to get the next model out
他们,我的意思是他们仍在追逐它,对吧,我确实认为他们仍在急于推出下一个模型
chasing it 常用搭配
追逐它、追求这个目标
chase 表示持续追求某个目标或梦想,常用于谈论理想、机会等。
rushing to get 常用搭配
急于推出、赶着完成
rush to do something 表示匆忙做某事,强调时间紧迫。
which will be much better than much as a right of term but will be better than the previous one
这将会比作为术语权利好得多,但会比前一个更好
and i i can't see them slowing down
而我,我看不到他们放慢脚步
slowing down 常用搭配
放慢速度、减速
slow down 可指速度变慢,也可指节奏或进展放缓。
i just think the gains will be made or felt more through not only scaling the model
我只是认为收益将更多地通过不仅扩展模型来实现或感受到
but now fine so i feel like there's a lot of tech depth
但现在好了,所以我觉得有很多技术深度
it's like well let's just put the better model in there and better
就像好吧,让我们把更好的模型放进去,然后更好
model better model and now people are okay let's also at the same time improve everything around it too like you know like the engineering of the context and inference scaling and i the big labs will still keep doing that and now also the smaller labs will catch up that because now it's just like they are hiring more there will be more people llms it's kind of like you know like a circle they also make them more productive and it's just it's like amplify i think what we can expect is amplification but not like a change of like a paradigm change i don't think that is true but everything will be just amplified and amplified and amplified and i can see that continuing for a long time you know yeah i guess my statement with the dream is dying depends on exactly what you think it's going to be doing like cloud code is a general model that can do a lot of things but it's not like necessarily like it depends a lot on integrations and other things like i bet cloud code could do a fairly good job of doing your email and the hardest part is figuring out how to
模型更好,模型更好,现在人们觉得可以了,让我们同时也改进它周围的一切,就像你知道的,比如上下文的工程和推理扩展,而大实验室仍然会继续做这些,现在小实验室也会赶上,因为现在就像他们在招更多人,会有更多人,LLMs,这有点像你知道的,像一个循环,他们也让他们更有效率,就像放大,我认为我们可以期待的是放大,但不是像范式改变那样的变化,我不认为那是真的,但一切都会被放大、放大、再放大,我可以看到这会持续很长时间,你知道,是的,我想我关于梦想正在消亡的说法取决于你认为它具体会做什么,比如Cloud Code是一个通用模型,可以做很多事情,但不一定像,这很大程度上取决于集成和其他事情,比如我敢打赌Cloud Code可以相当好地处理你的邮件,最难的部分是弄清楚如何
catch up 常用搭配
赶上、追上
catch up 表示落后之后追上来,常用于竞争或学习场景。
amplified and amplified 常用搭配
不断被放大、一再增强
重复同一个词表示程度不断加深,口语中常见这种强调方式。
give the information to it and how to get it to be able to send your emails and stuff like this
给它信息,以及如何让它能够发送你的邮件之类的东西
but that's just kind of like I think it goes back to
但这只是有点像,我认为这要追溯到
goes back to 常用搭配
回到、追溯到
go back to 表示话题或问题回到某个更根本的点上。
like what is the one model to rule everything
就像什么是那个统治一切的模型
ethos, which is just like a thing in the cloud that handles your entire digital life and is way smarter than everybody.
ethos,它就像云中的一个东西,处理你的整个数字生活,而且比所有人都聪明得多。
It's like it's operating in a...
就像它在一种……
So it's an interesting leap of faith to go from cloud code becomes that,
所以从云代码变成那样,是一个有趣的信仰之跃,
leap of faith 地道口语
信仰之跃、需要信念的一步
指在没有充分证据的情况下选择相信某事,常用于描述冒险的信任或推断。
which in some ways is there's some avenues for that.
这在某些方面,是有一些途径的。
But I do think that the rhetoric of the industry is a little bit different.
但我确实认为行业的说法有点不同。
I think the immediate also thing we will feel next as a normal person using LLMs will
我认为我们作为使用LLM的普通人接下来会感受到的直接的事情将
probably be related to something also trivial, like making figures.
可能也与一些同样琐碎的事情有关,比如制作图表。
Right now, LLMs are terrible at making figures.
现在,LLM在制作图表方面很糟糕。
terrible at 常用搭配
在……方面很糟糕
be terrible at something 表示非常不擅长某事,语气比 bad at 更强。
Is it because we are getting served the cheap models with very less, like, lesser inference compute than behind the scenes?
是因为我们得到的是廉价模型,其推理计算量比幕后少得多吗?
Maybe some, like, there are some cranks we can already get better figures.
也许有些,比如,有一些技巧我们已经可以获得更好的图表。
But if you ask today, I don't know, draw a flow chart of XYZ, it's most of the time terrible.
但如果你今天问,我不知道,画一个XYZ的流程图,大多数时候都很糟糕。
And it is kind of like a very simple task for a human.
而这对于人类来说是一个相当简单的任务。
I think it's almost easier sometimes to draw something than to write something.
我觉得有时候画点东西几乎比写点东西更容易。
Yeah, the multimodal understanding does feel like something that is odd
是啊,多模态理解确实感觉像是一件奇怪的事,
that it's not better solved.
它没有被更好地解决。
I think we're not saying one actually obvious thing
我觉得我们不是在说一个实际上很明显的事情,
that we're not actually realizing that's a gigantic thing that's hard to measure,
我们实际上没有意识到那是一件很难衡量的巨大事情,
which is making all of human knowledge accessible to the entire world.
那就是让全人类的知识对全世界都触手可及。
Like one of the things I think is hard to articulate,
就像我觉得很难说清楚的一件事,
but there's just a huge difference between Google search and an LLM.
但谷歌搜索和LLM之间有着巨大的区别。
Like I feel like I can basically ask an LLM anything and get an answer.
就像我觉得我基本上可以问LLM任何问题并得到答案。
less is doing less and less and less hallucination.
越少就是做得越少,幻觉也越来越少。
And that means understanding my own life,
而那意味着理解我自己的生活,
figuring out a career trajectory, figuring out how to solve the problems all around me,
弄清楚职业轨迹,弄清楚如何解决我周围的问题,
figuring out 常用搭配
弄清楚、搞明白
figure out 表示通过思考或调查弄明白某事,非常常用的口语表达。
learn about anything through human history.
通过人类历史学习任何东西。
I feel like nobody's really talking about that
我觉得没有人在真正谈论那个,
because they just immediately take it for granted that it's just, this is awesome.
因为他们立刻就想当然地认为,这只是,这太棒了。
take it for granted 地道口语
想当然、视为理所当然
take something for granted 表示认为某事理所当然而不加重视。
That's why everybody's using it is because you get answers for stuff.
这就是为什么大家都在用它,因为你能得到各种答案。
And like the impact of that across time, like think about this is not just in the United States.
而那随着时间推移的影响,想想看,这不仅仅是在美国。
It's all across the world.
它遍布全世界。
Like kids throughout the world being able to learn these ideas,
就像世界各地的孩子们能够学习这些理念,
like the impact that has across time is probably
就像它随着时间产生的影响,可能
that's where the real like talking about GDP.
这才是真正谈论GDP的地方。
It won't be like a leap.
它不会是一个飞跃。
It'll be that's how we get to Mars.
它会是——这就是我们到达火星的方式。
That's how we build these things.
这就是我们建造这些东西的方式。
That's why we have a million new open AIs, all the kind of innovation that happens from there.
这就是为什么我们有一百万个新的开放AI,以及由此产生的各种创新。
And that's just this quiet force that permeates everything, right?
而这只是一股渗透一切的安静力量,对吧?
Human knowledge.
人类知识。
I do agree with you.
我确实同意你的看法。
And in a sense, it makes knowledge more accessible.
在某种意义上,它让知识更容易获取。
But it also, I think, depends on what the topic is.
但我认为,这也取决于话题是什么。
For something like math, in a sense, you can ask it questions, it answers.
对于像数学这样的东西,在某种意义上,你可以问它问题,它来回答。
But if you want to learn a topic from scratch,
但如果你想从零开始学习一个话题,
from scratch 地道口语
从零开始、从头做起
表示完全从头开始做某事,没有任何基础或准备。
I think that, again, like we talked about this earlier, I think
我认为,就像我们之前谈到的,我认为
the sweet spot is, I mean, there are really good math textbooks where someone laid it out linearly.
最佳平衡点是,我的意思是,有些非常好的数学教科书,有人把它线性地编排出来。
And that is like, let's say, proven strategy to learn this topic.
那就像是,比如说,学习这个主题的经过验证的策略。
And it does make sense if you start from zero to ramp up to get like an information dense text to soak it up.
如果你从零开始逐步提升,去获取像信息密集的文本并吸收它,这确实是有道理的。
But then you use the LLM to make infinite exercises.
但然后你用LLM来制作无限的练习。
Like you have problems in a certain area or have questions and something's uncertain or like you're uncertain about certain things.
比如你在某个领域遇到问题,或者有疑问,或者某件事不确定,或者你对某些事情不确定。
You ask it to generate example problems.
你让它生成示例问题。
You solve them and you have questions.
你解决它们,然后你有疑问。
And then maybe you need more background knowledge and you ask it to generate that.
然后也许你需要更多背景知识,你就让它生成那些内容。
And I think then it won't give you anything, let's say, that is not in the textbook.
我觉得那样它就不会给你任何,比如说,不在教科书里的东西。
It's just packaging it differently, if that makes sense.
它只是换了一种方式包装,如果这么说有道理的话。
if that makes sense 地道口语
如果这么说有道理的话
说话人用来确认自己的表达是否被理解,常见于口语解释之后。
But then there are things I feel like where it also adds value in a more, I mean, timely sense where there is no good alternative besides a human doing it on the fly.
但有些情况我觉得它也能增加价值,在更——我是说——更及时的意义上,除了让人当场来做之外没有好的替代方案。
on the fly 地道口语
当场、临时即兴地
表示没有事先准备,在现场即时完成某事。
For example, if you, I don't know, like let's say you're planning to go to Disneyland and you try to figure out which tickets to buy for which park when, well,
比如说,如果你,我不知道,比如假设你打算去迪士尼乐园,你想弄清楚什么时候去哪个园区该买哪种票,那么,
there is no textbook on that.
关于这个没有教科书。
There is no information dense resource on that.
关于这个没有信息密集的资源。
There's only the sparse internet.
只有零散的互联网信息。
And then there is a lot of value in the LLM.
而这时大语言模型就有很大价值。
You just ask it.
你直接问它就行。
You have the constraints. I'm traveling these and these days.
你有这些限制条件。我这些天要去这里和那里。
I want to go there and there. Please figure out what I need, when and from where and what what it costs and stuff like that.
我想去那里和那里。请帮我弄清楚我需要什么、什么时候、从哪里出发,以及花费多少之类的事情。
And it is very customized on the fly, uh, package. And then this is
而且它是非常临时定制的,呃,套餐。然后这是
like one of the thousand examples. And exercise personalized, uh, personalization is essentially like pulling information from the sparse internet,
就像一千个例子中的一个。而练习个性化,呃,个性化本质上就像从稀疏的互联网中提取信息,
the non-information dense thing where there is no better version that exists. It just doesn't exist.
那种非信息密集的东西,那里不存在更好的版本。它就是不存在。
You make it from scratch almost.
你几乎是从零开始做。
from scratch 常用搭配
从零开始,白手起家
表示不借助已有的基础或材料,完全从头做起。
And if it does exist, it's full of, speaking of Disney World, like full of, what would you call it, ad slop?
而如果它确实存在,它充满了,说到迪士尼世界,就像充满了,你怎么称呼它,广告垃圾?
Like you just, it's impossible.
就像你只是,这不可能。
Here, you go any city in the world.
来,你去世界上任何一个城市。
What are the top 10 things to do?
排名前十的必做之事是什么?
LLM is just way better to ask than anything on the internet.
问大语言模型就是比问互联网上的任何东西都好得多。
Well, for now.
嗯,至少目前是这样。
That's because they're massively subsidized and they're going to be paid for by ads.
那是因为它们得到了大量补贴,而且将来会靠广告来支付。
It's coming.
这就要来了。
No.
不。
Oh, I hope there, I mean, I'm hoping there's a very clear indication of what's in that and what's not in that context.
哦,我希望,我是说,我希望有一个非常清晰的指示,说明那个语境中有什么、没有什么。
I did a little, you know, that's something I mentioned a few years ago. It's like...
我做过一点,你知道,那是我几年前提到过的事情。就像……
Uh, I don't know if you're looking for a new running shoe, Mom.
呃,我不知道你是不是在找新的跑鞋,妈妈。
Is this a coincidence that Nike maybe comes up first? Maybe, maybe not. And but I think there are clear laws around this. You have to be clear about that.
耐克可能最先出现,这是巧合吗?也许是,也许不是。但是我认为这周围有明确的法律。你必须对此保持清晰。
But I think that's what everyone fears. It's like the subtle, um, you know, subtle message in there or something like that.
但我认为那就是每个人所害怕的。就像那里面的微妙的,呃,你知道,微妙的信息或类似的东西。
But it also brings us to the topic of, I guess, ads.
但这也把我们带到了,我猜,广告的话题。
Where I think this was the thing, Open, may I try to launch in 2025.
我认为这就是那个东西,Open,我可以在2025年尝试推出。
And uh, just to, because I think it's still not uh making money in that other way right now. So that like having...
而且呃,只是因为我认为它现在还没有通过那种其他方式赚钱。所以那就像拥有……
Really like ad spots in there. And then the thing though is they couldn't because well there are alternatives.
真的喜欢在那里放广告位。但问题是他们不能,因为嗯,有替代品。
Without ads and people would just flock to the other products. And it also is just like crazy how...
没有广告,人们就会涌向其他产品。而且这也就像疯狂的是……
flock to 常用搭配
蜂拥而至,涌向
形容很多人迅速、大量地涌向某个地方或产品。
Yeah, like they're one-upping each other, spending so much money to just get the users.
是的,就像他们在互相攀比,花那么多钱只是为了获取用户。
one-upping each other 常用搭配
互相攀比,争相压过对方
形容双方不断比拼,试图比对方做得更好或更胜一筹。
I think so, like some Instagram ads. I don't use Instagram, but I understand the...
我想是的,比如一些Instagram广告。我不用Instagram,但我理解那种……
appeal of paying a platform to find users who will genuinely like your product and that is the best case of things like instagram ads but
付费给平台来找到真正喜欢你的产品的用户,这种吸引力,是像Instagram广告这样的最佳情况,但
there are also plenty of cases where advertising is very awful for incentives and i think that
也有很多情况下,广告对激励非常糟糕,我认为
a world where the power of ai can integrate with that positive view of like i am a person and i have a small business and i want to make the best i don't know damn steak knives in the world and i want to sell them to somebody who needs them and if like if ai can make that sort of advertising thing work even better that's very good for the world especially with like digital infrastructure because that's how like the modern web has been built but that's not to say like addicting feeds so that you can show people more content is a good thing so it's like i think that's even what opening i would say is they want to find a way that can make the monetization upside of ads while still giving their users agency and i'm i personally would think that google is probably going to be better at figuring out how to do this because they have
一个世界,在那里AI的力量可以与那种积极的视角相结合,比如我是一个人,我有一个小生意,我想做出世界上最好的,我不知道,该死的牛排刀,我想把它们卖给需要它们的人,如果AI能让那种广告方式运作得更好,那对世界非常有益,尤其是像数字基础设施,因为现代网络就是这样建立起来的,但这不是说像让人上瘾的信息流,这样你可以向人们展示更多内容,是件好事,所以就像我认为那甚至是开场我会说的,他们想找到一种方法,既能提高广告的变现收益,又能给用户自主权,而我个人会认为谷歌可能更擅长弄清楚如何做到这一点,因为他们已经
They already have ad supply and they figure out how to turn this demand
他们已经有广告供应,并且他们弄清楚如何将这种需求
in their Gemini app into useful ads
在他们Gemini应用中转化为有用的广告
then they can turn it on and somebody will figure
然后他们可以开启它,有人会弄清楚
I don't know if I think it's this year but
我不知道我是否认为这是今年,但是
there will be experiments with it I do think what holds companies back right now
会有关于它的实验,我确实认为现在阻碍公司的是
is really just that the competition is not doing it
真的只是竞争对手没有在做
it's more like more like a reputation thing it's just like
这更像是一种声誉问题,就像
I think people are just afraid right now like ruining or like losing the reputation losing users
我认为人们现在只是害怕,比如毁掉或失去声誉,失去用户
because it would make headlines if someone launched these ads.
因为如果有人推出这些广告,会成为头条新闻。
make headlines 常用搭配
成为头条新闻,引起广泛关注
形容某事因重大或引人注目而被媒体广泛报道。
Unless they were great.
除非它们很棒。
But the first ads won't be great because it's a hard problem that we don't know how to solve.
但最初的广告不会很棒,因为这是一个我们不知道如何解决的难题。
Yeah, I think also the first version of that will likely be something like on X, like the timeline where you have like a promoted post sometimes in between.
是的,我认为第一个版本可能就像在X上,比如时间线中有时会有一个推广帖子。
It will be something like that where it will say like promoted or something like small and then there will be an image or something.
它会像那样,会显示“推广”或类似的小字,然后会有一个图片之类的。
I think right now the problem is who makes the first move.
我认为现在的问题是,谁先迈出第一步。
makes the first move 常用搭配
先采取行动,先迈出第一步
指在竞争或谈判中率先行动的一方。
If we go 10 years out, the proposition for ads is that
如果我们展望10年后,广告的命题是
you will make so much money on ads by having so many users that you can use this to funnel
你会通过拥有如此多的用户,在广告上赚到如此多的钱,以至于你可以利用这一点来引导
better R&D and make better models, which is why like YouTube is dominating the market for any, like Netflix is scared of YouTube.
更好的研发并做出更好的模型,这就是为什么像YouTube正在主导任何市场,就像Netflix害怕YouTube一样。
Like they have the ad, like they make, I don't, I pay $28 a month for premium.
就像他们有广告,就像他们赚,我不,我每月付28美元买高级版。
They make at least $28 a month off of me and many other people.
他们每月至少从我以及许多其他人身上赚28美元。
And they're just like creating such a dominant position in video.
而他们就是在视频领域创造如此主导的地位。
So I think that's the proposition, which is that ads can make you have a sustained advantage in what you are spending per user
所以我认为这就是命题,即广告可以让你在每用户支出上拥有持续优势
but there's so much money in it right now that it's like like somebody starting that flywheel is scary
但现在这里面有太多钱了,以至于就像有人启动那个飞轮是可怕的
because it's a long-term bet uh do you think there'll be some like crazy big moves this year business-wise
因为这是一个长期赌注,呃,你觉得今年商业上会有一些疯狂的大动作吗
like somebody like google or apple acquiring anthropic or something like this dario will never sell but we are starting to see some types of consolidation with like grok for 20 billion dollars and um scale ai for almost 30 billion and countless other deals like this that they're structured in a way that is actually detrimental to
比如像谷歌或苹果收购Anthropic或类似的事情,达里奥永远不会卖,但我们开始看到一些类型的整合,比如Grok以200亿美元,以及呃Scale AI以近300亿美元,还有无数其他类似的交易,它们的结构方式实际上对……有害
The Silicon Valley ecosystem, which is this sort of licensing deal
硅谷生态系统,这是一种授权协议
where not everybody gets brought along, rather than a full acquisition
不是所有人都被一起带过去,而不是一次完整的收购
that benefits the rank and file employee by getting their stock vested like
这有利于普通员工,让他们的股票归属,就像
that's a big issue for Silicon Valley culture to address
这是硅谷文化需要解决的一个大问题
because the startup ecosystem is the lifeblood where
因为创业生态系统是命脉,在那里
if you get a, if you join a startup even if it's not that successful
如果你加入一家初创公司,即使它不那么成功
your startup very well might get acquired on a cheap premium of it and you'll get paid out for this equity and these licensing deals
你的初创公司很可能会以低价溢价被收购,你会因这些股权和这些授权协议获得报酬
are essentially taking the top talent a lot of the times
本质上往往是在获取顶尖人才
I think the deal for Grok to Nvidia is rumored to be better to the employees
我认为Grok与英伟达的交易据传对员工更好
but it is still this antitrust avoiding thing but I think
但它仍然是这种规避反垄断的做法,但我认为
that this trend of consolidation will continue.
这种整合趋势将会继续。
I've been, me and many smart people I respect have been expecting consolidation to have happened sooner,
我一直,我和许多我尊敬的聪明人一直预期整合会更早发生,
but it seems like some of these things are starting to turn, which,
但看起来其中一些事情开始转变,这,
but at the same time you have companies raising ridiculous amounts of money for reasons that you don't, like, I'm like, I don't know why you're taking that money.
但与此同时,有些公司筹集了巨额资金,原因你并不,就像,我想,我不知道你为什么要拿那笔钱。
So it's maybe like mixed this year, but some consolidation pressure is starting.
所以今年可能算是喜忧参半,但一些整合压力正在开始。
What kind of surprising consolidation do you think we'll see?
你觉得我们会看到什么样的出人意料的整合?
So you're saying Anthropic is a never.
所以你是说Anthropic永远不会。
I mean, Grok is a big one. Grok with a Q, by the way.
我是说,Grok是个大热门。顺便说一句,Grok是Q结尾的。
Yeah. There's just a lot of startups, and there's a very high premium on AI startups.
是的。现在有很多初创公司,而且AI初创公司的估值溢价非常高。
So there's a lot of like, there could be a lot of 10 billion range acquisitions,
所以会有很多,可能会有很多百亿美元级别的收购,
which is a really big acquisition for a startup that was maybe founded like a year ago.
这对一家可能一年前才成立的初创公司来说,是一笔非常大的收购。
I think Manus AI, the company that's based in Singapore that Meta founded, was founded eight months ago and then had a $2 billion exit.
我觉得Manus AI,那家总部在新加坡、Meta创立的公司,成立才八个月,然后就以20亿美元退出。
And I think that there will be some other big, like many billion dollar acquisitions, like perplexity.
而且我觉得还会有其他大额的、几十亿美元级别的收购,比如Perplexity。
Yeah, people rumored them to Apple.
是的,有人传言它们会被苹果收购。
I think there's a lot of pressure and liquidity in AI.
我觉得AI领域有很多压力和流动性。
There's pressure on big companies to have outcomes.
大公司面临着要出成果的压力。
And I would guess that a big acquisition gives people leeway to then tell the next chapter of that story.
我猜一笔大收购会给人余地,去讲述那个故事的下一章。
I mean, yeah, I guess cursor.
我是说,是的,我猜是Cursor。
We've been talking about code and somebody acquires cursor.
我们一直在聊代码,然后有人收购了Cursor。
They're in such a good position by having so much user data.
他们因为拥有这么多用户数据,处于非常有利的位置。
in such a good position 常用搭配
处于非常有利的处境
用来形容某人或某公司因具备某种优势而占据有利地位。
Yeah. And we talked about continual learning and stuff.
是的。我们还聊到了持续学习之类的。
and stuff 地道口语
之类的、等等
口语中放在列举之后,表示还有其他类似的东西,不必一一说明。
They had one of the most interesting, like, two sentences in a blog post,
他们在一篇博客文章里写了最有趣的,大概两句话,
which is that they had their new composer model,
就是说他们有了新的作曲模型,
which was a fine tune of one of these large mixture of expert models from China.
这是对中国某个大型专家混合模型进行微调得到的。
You can know that by asking gossip
你可以通过问八卦来知道这一点,
or because the model sometimes responds in Chinese,
或者因为这个模型有时会用中文回答,
which none of the American models do.
而美国的模型都不会这样。
And they had a blog post where they're like,
他们还有一篇博客文章,里面说,
we're updating the model weights every 90 minutes
我们每90分钟更新一次模型权重,
based on real world feedback from people using it,
根据使用它的人的真实世界反馈,
which is like the closest thing to real world RL happening on a model.
这就像是在模型上发生的最接近真实世界强化学习的事情。
the closest thing to 句型
最接近……的东西
the closest thing to [something]
用来形容某事物虽然不是完全等同,但已是最接近某种情况的例子。
And it's just like in one of their blog posts, which is super cool.
这就像在他们的一篇博客文章里说的,超级酷。
And by the way, I should say I use Composer a lot
顺便说一句,我应该说我经常用Composer,
by the way 地道口语
顺便说一句
口语中用来插入一个与当前话题相关但非重点的补充信息。
because one of the benefits it has is it's fast.
因为它有一个好处就是快。
I need to try it because everybody says this.
我需要试试它,因为大家都这么说。
And there'll be some IPOs potentially.
而且可能会有一些IPO。
You think Anthropic, OpenAI, XAI?
你觉得Anthropic、OpenAI、XAI会吗?
They can all raise so much money so easily that they don't feel a need to.
他们都能如此轻易地筹集到这么多钱,以至于他们觉得没必要上市。
so much money so easily 常用搭配
如此轻易地筹到这么多钱
用来强调融资既容易又数额巨大,常带感叹或对比语气。
Like so long as fundraising is easy, they're not going to IPO
就像只要融资容易,他们就不会IPO,
so long as 常用搭配
只要
引导条件从句,表示在某条件持续成立的情况下,主句结果就成立。
because public market supply pressure.
因为公开市场的供应压力。
I think we're seeing in China that the ecosystem's a little different
我认为我们在中国看到生态系统有点不同,
with both Minimax and Z.ai applying for filing IPO paperwork,
Minimax和Z.ai都在申请提交IPO文件,
which will be interesting to see how the Chinese market reacts.
看看中国市场会如何反应会很有趣。
I actually would guess that it's going to be like similarly hypey to the US, so
我其实猜它会像美国一样炒作得很热,所以
long as all this is going and not based on the realities that they're both losing a ton of money.
只要这一切还在继续,而不是基于它们都在亏大钱的现实。
losing a ton of money 常用搭配
亏大钱
口语中形容亏损数额巨大,语气比 lose money 更强烈。
I wish more of the American gigantic AI startups were public because it would
我希望更多美国的大型AI初创公司是上市的,因为那样会
I wish more of 句型
我希望更多……
I wish more of [something] [were/did something]
用 wish 表达对现状不满、希望情况不同的愿望,后接名词或从句。
be very interesting to see how they're spending their money and have more insight.
很有意思,能看到它们怎么花钱,也能有更多了解。
And also just to give people access to investing in these because I think that they're some of the most like formative like they're the companies of the era and the tradition is now for so many of the big startups in the U.S. to not go public
而且也是为了让人们能投资这些公司,因为我觉得它们是最具塑造性的,像是这个时代的公司,而现在传统上美国很多大型初创公司都不上市
it's like we're still waiting for Stripe and the IPO but Databricks definitely didn't they raised like a Series G or something and
就像我们还在等Stripe和它的IPO,但Databricks肯定没有,它们融了G轮之类的,而且
I just feel like it's a kind of a weird equilibrium for the market where it's like I would like to see these companies go public and evolve in that way that a company can, you think
我就觉得市场处于一种奇怪的平衡,就像我希望看到这些公司上市,并以公司能够发展的方式演进,你觉得
10 years from now some of the frontier model companies are still around, Anthropic, OpenAI, I definitely don't see it to be a winner takes all unless there truly is so algorithmic secret that one of them finds like let's say flywheel because the development path is
十年后一些前沿模型公司还在,Anthropic、OpenAI,我绝对不认为会是赢家通吃,除非真的存在某种算法秘密,被其中一家找到,比如飞轮效应,因为发展路径是
winner takes all 常用搭配
赢家通吃
形容竞争中最终只有一个赢家独占全部市场或收益的局面。
So similar for all of them, Google and OpenAI have like all the same products, and then like Anthropic's more focused,
所以它们都差不多,谷歌和OpenAI有几乎一样的产品,而Anthropic更专注,
but when you talk to people, it sounds like they're solving a lot of the same problems.
但当你和人们交谈时,听起来他们在解决很多相同的问题。
So I think, and there's offerings that'll spread out.
所以我认为,而且会有各种产品分散开来。
There's a lot of, it's a very big cake that's being made that people are going to take money out of.
有很多,这是一个正在做大的蛋糕,人们会从中赚钱。
I don't want to trivialize it, but so OpenAI and Anthropic are primarily LLM service providers.
我不想轻视它,但OpenAI和Anthropic主要是LLM服务提供商。
And some of the other companies like Google and XAI, linked to X, does other stuff too.
而其他一些公司,比如谷歌和XAI(与X关联),也做其他事情。
And so it's very possible if AI becomes more commodified, that the companies that are just providing LLM will die, I think.
所以很有可能,如果AI变得更加商品化,那些仅仅提供LLM的公司会消亡,我认为。
They will, the advantage they have, they have a lot of users, and I think they will just pivot, I think, um,
它们会,它们拥有的优势,它们有很多用户,我认为它们会转型,我认为,嗯,
then, uh, if they figure out it's like Anthropic, I think pivoted, I don't think they originally, um, planned to work on code,
然后,呃,如果他们发现像Anthropic那样,我认为转型了,我不认为他们最初,嗯,计划做代码,
but it happened that they found, okay, this is like a nice niche, and now we are comfortable in this niche, and we push on this niche, and I can see the same thing once maybe,
但碰巧他们发现,好吧,这是一个不错的细分领域,现在我们在这个领域很自在,我们在这个领域推进,我能看到同样的事情,也许一旦,
let's say hypothetically speaking, I'm not sure it will be true, but let's say Google takes all
假设来说,我不确定这是否会成真,但假设谷歌占据所有
hypothetically speaking 常用搭配
假设来说
用于引出一种假设性的、未必真实的情况,提醒对方这只是设想。
The market share of the general chatbot, maybe OpenAI, I will be then focus on some other subtopic like
通用聊天机器人的市场份额,也许是 OpenAI,我将会专注于其他一些子话题,比如
the they have too many users to go away in the foreseeable future
他们有太多用户,在可预见的未来不会消失
in the foreseeable future 常用搭配
在可预见的未来
表示在可以预料的未来一段时间内,常用于预测或判断。
I think I think
我认为我认为
Google is always ready to say hold my beer with AI mode
谷歌总是准备好说“看我的”,推出 AI 模式
hold my beer 地道口语
看我的、让我来露一手
网络和口语中表示某人准备展示更强实力、超越前面被夸赞的对象,常带幽默或挑衅意味。
I think that the question is if the companies can support the valuations
我认为问题在于这些公司能否支撑这些估值
I think I'd see the AI companies being looked at in some ways like AWS, Azure and GCP are all competing in the same space and all very successful businesses
我认为我会看到 AI 公司在某种程度上被看作像 AWS、Azure 和 GCP 一样,它们都在同一领域竞争,而且都是非常成功的企业
There's a chance that the API market is so unprofitable that they go up and down the stack to products and hardware
有可能 API 市场太不赚钱,以至于他们会在技术栈上下游扩展到产品和硬件
They have so much cash that they can build power plants and build data centers, which is a durable advantage now
他们有如此多的现金,可以建造发电厂和数据中心,这在现在是一种持久的优势
But there's also just a reasonable outcome that these APIs are so valuable and so flexible for developers that they become the likes of like something like AWS, but AWS and nature are also going to have these APIs, so there's some like that's a like five or six people competing in the API market is hard, so maybe like that's why they get squeezed out
但也有一个合理的结果,即这些 API 对开发者来说如此有价值且灵活,以至于它们变得像 AWS 一样,但 AWS 和自然也会拥有这些 API,所以有些像那样,五六个人在 API 市场竞争很难,所以也许这就是它们被挤出的原因
You mentioned RIP Llama, is there a path to winning for Meta? I think nobody knows
你提到 RIP Llama,Meta 有获胜的路径吗?我认为没人知道
They're moving a lot, so they're signing licensing deals with, um, Black Forest Labs, which is an image generation or Midjourney or requiring mainness.
他们动作很大,所以他们在和,嗯,黑森林实验室签授权协议,那是一个图像生成,或者Midjourney,或者需要主流化。
So I think in some ways it's on the product and like consumer facing AI front.
所以我认为在某些方面,它是在产品和面向消费者的AI前沿上。
It's too early to tell. I think they have some people that are excellent and very motivated being close to Zuckerberg.
现在下结论还为时过早。我认为他们有一些优秀且非常积极的人,与扎克伯格关系密切。
too early to tell 常用搭配
现在下结论还为时过早
表示目前信息不足,无法判断结果或趋势,需要再观察。
So I think that there's still a story to unfold there.
所以我认为那里仍然有一个故事要展开。
Llama is a bit different, where Llama was the most focused expression of the organization, and I don't see Llama being supported to that extent.
Llama有点不同,Llama是该组织最专注的表达,而我看不到Llama会得到那种程度的支持。
I think it was a very successful brand for them, so they still might do some part of participation in the open ecosystem or continue the Llama brand into a different surface,
我认为这对他们来说是一个非常成功的品牌,所以他们仍然可能参与开放生态系统的一部分,或者将Llama品牌延续到一个不同的界面,
because people know what Llama is. You think there's a Llama 5, not an open weight one? It's interesting.
因为人们知道Llama是什么。你认为会有Llama 5,但不是开放权重的那种?这很有趣。
I think also just to recap a bit, I think I mean Llama was the, I would say, pioneering open weight model, and then Llama 1, 2, 3 a lot of love, but I think then I think what happened, just hypothesizing or speculating, I think the
我还想简单回顾一下,我认为Llama是,我会说,开创性的开放权重模型,然后Llama 1、2、3获得了很多喜爱,但我想然后我想发生了什么,只是假设或推测,我认为
um, leaders at Meta, like the upper, uh, executives, they, I think they got really excited about Llama because they saw how popular it was in the community.
嗯,Meta的领导者,比如高层,呃,高管们,我认为他们对Llama非常兴奋,因为他们看到它在社区中多么受欢迎。
And then I think the problem was trying to...
然后我觉得问题在于试图……
let's say monetize the open, not monetize the open source,
比如说将开源变现,不是将开源变现,
but like kind of use the open source to make a bigger splash in a sense,
而是有点像利用开源来制造更大的声势,
make a bigger splash 常用搭配
制造更大的声势、引起更大轰动
形容某事引起广泛关注或强烈反响,常与 in a sense 等搭配。
like to kind of force it almost, it felt forced like...
有点像几乎强迫它,感觉是被迫的,就像……
developing these very big Llama 4 models to have like the best, like to be on the top of the benchmarks,
开发这些非常大的Llama 4模型,以拥有最好的,比如在基准测试中名列前茅,
but I don't think the goal of Llama models is to be on top of the benchmarks,
但我不认为Llama模型的目标是在基准测试中名列前茅,
beating, let's say, Chatopia or other models. I think the goal was to have a model that people can use, trust, modify, understand. That so that includes
击败,比如说,Chatopia或其他模型。我认为目标是拥有一个人们可以使用、信任、修改、理解的模型。所以这包括
having smaller models. They don't have to be the best models, and what happened
拥有更小的模型。它们不必是最好的模型,而发生的事情
was just these models were, of course, like the benchmarks suggested they were better than they were, because I think they had like specific models trained on preferences that they perform well on the benchmarks.
只是这些模型当然就像基准测试所显示的那样比实际更好,因为我认为他们有特定的模型在偏好上训练,使它们在基准测试中表现良好。
So it's kind of like this overfitting thing to kind of force it to be the best, but then at the same time
所以这有点像这种过拟合的事情,有点像强迫它成为最好的,但与此同时
they didn't do the small models that people could use, anything that no one could run these big models then. And then there was kind of like a weird thing, and I think it's just because people
他们没有做人们可以使用的小模型,任何没有人能运行这些大模型的东西。然后就有一种奇怪的事情,我认为这只是因为人们
got too excited about headlines pushing the frontier
对推动前沿的头条新闻太兴奋了
I think, and too much like on the benchmarking side
我觉得,而且在基准测试这方面也太多了
Yeah, I think it imploded under political, like, internal political fighting and misaligned incentives
是的,我觉得它因为政治,像是内部政治斗争和错位的激励而崩溃了
So I think the researchers want to build the best
所以我觉得研究人员想打造最好的
models but there's a layer of organization and manager that is trying to demonstrate
模型,但有一层组织和管理者试图证明
that they do these things and then
他们做了这些事情,然后
there's lots of, there's a lot of pieces and rumors where
有很多,有很多碎片和传闻,说
how, like, some horrible technical decision was made and how that comes in.
怎么,像是,做出了某个糟糕的技术决策,以及那是怎么进来的。
And it just seems like it kind of got too bad where it all just crashed out.
而且看起来它变得太糟了,最后全都崩盘了。
crashed out 地道口语
彻底崩溃、垮掉
口语中形容系统、项目或局面彻底失败、崩盘。
We should also give huge props to Mark Zuckerberg.
我们也应该给马克·扎克伯格大大的赞赏。
give huge props to 常用搭配
给予高度赞赏
口语中表示公开称赞或致敬某人,props 即 respect。
I think it comes from Mark, actually.
我觉得这其实来自马克。
From Mark Zuckerberg, from the top of the leadership, saying open source is important.
来自马克·扎克伯格,来自最高领导层,说开源很重要。
I think that's like, the fact that that exists means there could be a Lama 5
我觉得这就像是,它存在这个事实意味着可能会有个Lama 5
where they learn the lessons from the bench maxing and say
在那里他们从基准刷分中吸取教训,然后说
we're going to be gpt oss and provide really awesome library of open source
我们要做gpt oss,并提供非常棒的开放源代码库
what people say is that there's a debate between mark and alexander wang
人们说的是,马克和亚历山大·王之间有场争论
who is very bright but much more against open source
他非常聪明,但更反对开源
and to the extent that he has a lot of influence over the ai org it seems much less likely
而就他对AI组织有很大影响力而言,这似乎就不太可能了
to the extent that 句型
就……而言;在……的程度上
to the extent that [condition], [consequence]
用于表示某条件成立到什么程度,后接结果或推论。
because it seems like mark brought him in
因为看起来是马克把他招进来的
brought him in 常用搭配
把他招进来、引入团队
指把某人引进组织或项目,常与 for/to do 搭配。
for like a fresh leadership aid in directing AI.
为了像是新的领导层来协助指导AI。
And if the like open or closed is no longer the defining nature of the model,
如果开放或封闭不再是模型的本质定义,
I don't expect that to be a defining argument between Mark and Alex.
我不认为那会成为马克和亚历克斯之间的决定性争论。
Like they're both very bright, but I just like, I have a hard time understanding all of it
他们俩都很聪明,但我就是,我很难理解这一切,
I have a hard time understanding 句型
我很难理解……
have a hard time [doing something]
表达做某事有困难,后接动名词,语气自然。
because Mark wrote this piece in July of 2024, maybe,
因为马克在2024年7月左右写了这篇文章,
which was like probably the best blog post at the time saying the case for open source AI.
那大概是当时最好的博客文章,阐述了开源AI的理由。
And then July 2025 came around and it was like, we're reevaluating our relationship with open source.
然后到了2025年7月,就好像,我们在重新评估与开源的关系。
So it's just kind of like.
所以就是有点像是。
But I think also the problem, not the problem, but I think, well, we may have been a bit also too harsh, I think.
但我觉得问题,不是问题,但我觉得,嗯,我们可能也有点太严厉了,我觉得。
And that caused some of that because I think, I mean, we as open source developers or the open source community,
而那导致了一些这样的情况,因为我觉得,我是说,我们作为开源开发者或开源社区,
because I think even though the model was maybe not what everyone hoped for, it got a lot of backlash.
因为我觉得即使这个模型可能不是大家所期望的,它也遭到了很多反对。
got a lot of backlash 常用搭配
遭到大量反对/抨击
形容因某言行或产品引发强烈负面反应。
And I think that was a bit unfortunate because I can see that as a company.
我觉得那有点不幸,因为作为一家公司我能理解。
Now they were hoping for positive headlines and instead of just getting no headlines or not these positive headlines, in turn, they got negative headlines.
他们本来希望得到正面报道,结果不仅没有报道或者不是这些正面报道,反而得到了负面报道。
And then all it kind of reflected bad on the company.
然后这一切都对公司造成了不好的影响。
reflected bad on 常用搭配
给……带来不好的影响
指某行为或事件使某人/公司形象受损,常用 reflect badly on。
And I think that is also something like where you, it's maybe a spite reaction,
而且我觉得那也像是,你,这也许是一种逆反反应,
almost like, okay, we have, we tried to do something nice.
几乎像是,好吧,我们有过,我们试图做点好事。
We try to give you something cool, like an open source model.
我们试图给你一些酷的东西,比如一个开源模型。
And now you are like, you know, kind of like being negative about us, even like for the company.
而现在你就像,你知道,有点对我们持负面态度,甚至对公司也是如此。
So in that sense, it looks like, well, maybe then we'll change our mind, I guess.
所以从这个意义上说,看起来,好吧,也许那我们会改变主意,我猜。
I don't know.
我不知道。
Yeah, that's where the dynamics of discourse on X can lead us as a community astray.
是的,这就是X上的话语动态可能把我们作为一个社区引入歧途的地方。
lead us as a community astray 常用搭配
把我们这个群体引入歧途
lead someone astray 指使人误入歧途或做出错误判断。
Because sometimes it feels random.
因为有时候感觉是随机的。
People pick the thing they like, they don't like.
人们选择他们喜欢的东西,他们不喜欢。
And you can see the same thing with Grok4.1 and GrokCodeFast1.
你可以看到Grok4.1和GrokCodeFast1也有同样的情况。
I don't think vibe-wise people love it publicly,
我不认为从氛围上来说人们公开喜欢它,
but a lot of people use it.
但很多人使用它。
So if you look to Reddit and X, they don't really give it praise from the programming community,
所以如果你看看Reddit和X,编程社区并不真正赞扬它,
but they use it.
但他们使用它。
And the same thing with probably with Lama.
可能Lama也有同样的情况。
I don't understand the dynamics of either positive hype or negative hype.
我不理解正面炒作或负面炒作的动态。
I don't understand it.
我不理解。
I mean, the story of one of the stories of 2025 is the U.S.
我的意思是,2025年的故事之一是美国。
feeling the gap of llama, which is like all the rise of these Chinese open-way models to the point
感受到 Llama 的差距,就像所有这些中国开放权重模型的崛起,到了那个地步
where I was like, that was the single issue
我当时觉得,那就是唯一的问题
I've spent a lot of energy on in the last five months
我在过去五个月里投入了大量精力
spent a lot of energy on 常用搭配
在……上投入了大量精力
表示把时间精力花在某事上,后接名词或动名词。
is like trying to do policy work to get the U.S. to invest in this.
就是试图做政策工作,让美国投资于这个领域。
Tell me the story of Adam.
告诉我亚当的故事。
Atom Project, it started as me calling it the American Deep Seek Project,
Atom 项目,一开始我把它叫做美国深度求索项目,
which doesn't really work for DC audiences,
这个名字对华盛顿的受众不太管用,
but it's the story of what is the most impactful thing I can do with my career,
但这是关于我职业生涯中最有影响力的事情的故事,
which is that Chinese open weight models are cultivating a lot of power and there is a lot of demand for building on these open models,
那就是中国开放权重模型正在积聚大量力量,而且有很多需求要在这些开放模型上构建,
especially in enterprises in the US that are very cagery about these Chinese models.
尤其是在美国的企业中,它们对这些中国模型非常谨慎。
Go to perplexity.
去用 Perplexity。
The Atom Project, American Truly Open Models, is a U.S.-based initiative to build and host high-quality, genuinely open-weight AI models and supporting infrastructure explicitly aimed at competing with and catching up to China's rapidly advancing open-source AI ecosystem.
Atom 项目,即美国真正开放模型,是一个总部位于美国的倡议,旨在构建和托管高质量、真正开放权重的 AI 模型及支持性基础设施,明确目标是与快速发展的中国开源 AI 生态系统竞争并迎头赶上。
I think the one sentence summary would be that, or two sentences,
我认为一句话的总结会是,或者两句话,
one is a proposition that open models are going to be an engine for AI research
一是这样一个命题:开放模型将成为AI研究的引擎,
because that is what people start with.
因为那是人们起步时所用的东西。
Therefore, it's important to own them.
因此,拥有它们很重要。
And the second one is, therefore, the U.S. should be building the best models so that the best research happens in the U.S. and the U.S. companies take the value from being the home of where AI research is happening.
第二点是,因此,美国应该打造最好的模型,以便最好的研究在美国发生,美国公司从作为AI研究发生地的所在地中获取价值。
And without more investment in open models,
如果没有对开放模型的更多投资,
we have all the plots on the website where it's like, Quinn, Quinn, Quinn, Quinn. And it's all these models that are excellent from these Chinese companies that are cultivating influence in the U.S. and China and internationally.
我们在网站上展示了所有图表,上面就像,Quinn、Quinn、Quinn、Quinn。全都是这些来自中国公司的优秀模型,它们正在美国、中国以及国际上培养影响力。
And I think the U.S. is spending way more on AI and the ability to create open models that are half a generation or a generation beyond what the cutting edge of a closed labs is costs orders of like $100 million, which is a lot of money, but not a lot of the money to these companies.
而且我认为美国在AI上的支出要多得多,而能够创造出比闭源实验室最前沿还要领先半代或一代的开放模型,其成本大约在1亿美元左右,这是一大笔钱,但对这些公司来说并不算多。
So therefore, we need a centralizing force of people who want to do this.
所以,我们需要一股想要做这件事的人们的集中力量。
And I think.
而且我认为。
We got signed engagement from people pretty much across the full stack, whether it's policy.
我们得到了几乎整个技术栈上的人签署的承诺,无论是政策方面。
So there has been support from the administration?
所以政府方面一直有支持吗?
I don't think anyone in that, like technically in government has signed it publicly,
我不认为政府里有人,从技术上讲,公开签署了它,
but I know that people that have worked in AI policy,
但我知道那些从事过AI政策工作的人,
both in Biden and Trump administration are very supportive of trying to promote open source models in the US.
在拜登和特朗普政府中,都非常支持在美国推广开源模型。
I think, for example, AI2 got a grant from the NSF for a hundred million dollars over four years,
我认为,例如,AI2从NSF获得了四年一亿美元的资助,
which is like the biggest cs grant the nsf has ever awarded
这就像是NSF有史以来颁发的最大一笔计算机科学资助
and it's for ai2 to attempt to this
而且这是为了让AI2尝试做这个
and i think it's a starting point
我认为这是一个起点
but the best thing happens when there are multiple organizations building models
但最好的情况是当有多个组织在构建模型时
because they can cross-pollinate ideas and kind of build this ecosystem
因为他们可以交流想法并建立这个生态系统
cross-pollinate ideas 常用搭配
交流、碰撞想法
比喻不同团队或领域之间互相启发、交换创意。
like i don't think if it just works if it's just llama releasing models to the world
就像我不认为如果只是Llama向世界发布模型就能行得通
because then you can see llama can go away the same thing applies for ai2
因为那样的话,你可以看到Llama可能会消失,同样的情况也适用于AI2
where it's like i can't be the only one building models
就像我不能是唯一一个构建模型的人
and i think that's like that it becomes a lot of time spent on talking to people
我认为这就像变成了大量时间花在与人们交谈上
spent on 常用搭配
花在……上
用于描述时间、金钱或精力被用于某事物,常见搭配 time/money spent on doing something。
whether they're in policy i know nvidia is very excited about this i think
无论他们是否在政策领域,我知道英伟达对此非常兴奋,我认为
Jensen Wong has been specifically talking about the urgency for this and
Jensen Wong一直在特别谈论这件事的紧迫性,而且
talking about the urgency for 常用搭配
谈论……的紧迫性
用于强调某件事需要尽快处理,常与 urgency for/of 搭配。
they've changed, they've done a lot more in 2025
他们变了,他们在2025年做了更多
where the Nematron models are more of a focus
Nematron模型更受关注
they've started releasing some data along with Nvidia's open models
他们已经开始发布一些数据,还有Nvidia的开放模型
along with 常用搭配
连同……一起
表示附加或伴随关系,比 and 更强调一并出现。
and like very few companies do this, especially of Nvidia's size
而且很少有公司这样做,尤其是像Nvidia这样规模的
so like there is, there is signs of progress and
所以就像,有进步的迹象,而且
signs of progress 常用搭配
进步的迹象
用于描述某领域出现好转或向前发展的苗头。
there we hear about Reflection AI
我们听说Reflection AI
where they say their two billion dollar fundraise is dedicated to building U.S. open models
他们说他们的20亿美元融资专门用于构建美国开放模型
dedicated to 常用搭配
专门用于
表示资金、时间或努力被专门投入某个目的。
and I feel like their announcement tweet is like it reads like a blog post out right
我觉得他们的公告推文读起来就像一篇博客文章,对吧
reads like 常用搭配
读起来像
用于评价文字给人的感觉,常指某文本风格像另一类文体。
and I think that that cultural tide is starting to turn
我认为那种文化潮流开始转变了
I think in, in July was when we had like four or five deep sea caliber
我认为在,在七月的时候,我们有四五个深海级别的
Chinese open weight models in zero from the US.
中国开放权重模型,来自美国的零个。
And that's the moment where it was released this.
那就是它发布这个的时刻。
And I was like, oh, I guess I have to spend energy on this
我当时想,哦,我想我得在这上面花精力
I guess I have to 句型
我想我得……
I guess I have to [do something]
口语中表示不太情愿但不得不做某事。
because nobody else is going to do it.
因为没有其他人会去做。
So it takes a lot of people contributing together.
所以需要很多人一起贡献。
And I don't say that like the Atom project isn't like the thing that's helping to move the ecosystem,
我不是说Atom项目不是帮助推动生态系统的东西,
but it's people like me doing this sort of thing to get the word out.
但正是像我这样的人做这种事来传播信息。
get the word out 常用搭配
把消息传播出去
指让更多人知道某件事,常用于宣传、推广或告知信息。
Do you like the 2025 America's AI action plan that includes open source stuff?
你喜欢2025年美国人工智能行动计划中包括开源内容的部分吗?
The White House AI action plan includes a dedicated section titled to encourage open source and open web AI, defining such models and arguing they have unique value for innovation and startups.
白宫人工智能行动计划包括一个专门章节,旨在鼓励开源和开放网络人工智能,定义此类模型并主张它们对创新和初创企业具有独特价值。
Yeah. I mean, like the AI action plan is a plan, but largely I think it's like maybe the most coherent policy document that has come out of the administration.
是的。我的意思是,人工智能行动计划是一个计划,但很大程度上我认为它可能是政府出台的最连贯的政策文件。
And I hope that it largely succeeds.
我希望它大体上能成功。
And I know people that have worked on the AI action plan and the challenge is taking policy and making it real.
我认识参与人工智能行动计划的人,挑战在于把政策变为现实。
making it real 常用搭配
把它变为现实
指把计划、政策或想法真正落实。
And I have no idea how to do this as an AI researcher.
作为一名人工智能研究员,我完全不知道如何做到这一点。
I have no idea how to 句型
我完全不知道如何……
I have no idea how to [do something]
口语中表示对做某事毫无头绪。
But like, I think largely a lot of things in that were very real.
但就像,我认为其中很多事情是非常真实的。
And there's a huge build out of AI in the country.
而且这个国家正在进行大规模的人工智能建设。
And it's like, there are a lot of issues that people are hearing about from water use to whatever.
就像,人们听说有很多问题,从用水到其他各种事情。
And like, we should be able to build things in this country.
而且就像,我们应该能够在这个国家建设东西。
But also we need to not ruin places in our country in the process of building it.
但我们也需要在建设过程中不破坏我们国家的地方。
And it's worthwhile to spend energy on.
这是值得投入精力的。
worthwhile to spend energy on 常用搭配
值得投入精力
用于说明某件事值得花时间和精力去做。
I think that's a role that the federal government plays.
我认为这是联邦政府扮演的角色。
It's like they set the agenda.
就像他们设定了议程。
set the agenda 常用搭配
设定议程
指决定讨论或行动的重点方向,常用于政策、会议或组织语境。
And with AI setting the agenda, that open weight should be a first consideration,
而由AI设定议程,开放权重应该是首要考虑因素,
that's a large part of what they can do.
这是他们能做的很大一部分。
And then people think about it.
然后人们会思考它。
Also for education and talent for these companies, it's, I think, very important
此外,对于这些公司的教育和人才来说,我认为非常重要,
because otherwise, you know, if they're only closed models,
因为否则,你知道,如果它们只是闭源模型,
how do you get the next generation of people contributing at some point?
你如何在某个时候让下一代人做出贡献?
Because otherwise you will at some point only be able to learn after you joined a company.
因为否则你在某个时候只能在加入公司后才能学习。
But then at that point, like how do you hire talented people?
但到那时,就像你如何雇佣有才华的人?
How do you identify talented people?
你如何识别有才华的人?
And I think open source is, that's even for a lot of things, but also even just for educating the population and training the next generation of researchers.
我认为开源是,这不仅对很多事情,甚至只是为了教育大众和培训下一代研究人员。
It's the way or the only way.
这是方式,或者唯一的方式。
The way that I could have gotten this to go more viral was to tell a story of Chinese AI integrating with an authoritarian state and being ASI and taking over the world.
我本可以让这件事更病毒式传播的方式是讲述一个中国AI与威权国家结合、成为ASI并接管世界的故事。
And therefore, we need our own American models.
因此,我们需要自己的美国模型。
But it's very intentional for why I talk about innovation and science in the US,
但我谈论美国的创新和科学是非常有意的,
because I think it's both more realistic as an outcome,
因为我认为这作为结果更现实,
but just like it's a world that I would like to manifest.
但就像这是一个我想要实现的世界。
I would say, though, also even like, let's say, any open-weight model I do think is a valuable model.
不过,我想说,甚至比如,任何开放权重的模型,我确实认为都是有价值的模型。
Yeah, my argument is that we should be in a leading position.
是的,我的论点是,我们应该处于领先地位。
But I think that it's worth saying it so simply because there are still voices in the AI ecosystem that say we should consider banning releasing open models due to the safety risks.
但我认为值得简单地说,因为AI生态系统中仍然有声音说,出于安全风险,我们应该考虑禁止发布开放模型。
And I think it's worth adding that I think effectively that's impossible without making the U.S. have its own great firewall,
我认为值得补充的是,我认为如果不让美国拥有自己的防火墙,这实际上是不可能的,
which is also known to not work that well
而众所周知,防火墙的效果并不那么好,
because the cost for training these models, whether it's one to a hundred million dollars, is attainable to a huge amount of people in the world that want to have influence.
因为训练这些模型的成本,无论是100万到1亿美元,对于世界上大量想要拥有影响力的人来说都是可以承受的。
So these models will be getting trained all over the world.
所以这些模型将在世界各地接受训练。
And these we want the models, especially like I mean, there are safety concerns, but we want these information and tools to flow freely across the world and into the U.S. so that we people can use them and learn from them.
而我们希望这些模型,特别是,我的意思是,存在安全担忧,但我们希望这些信息和工具能够自由地跨越世界并进入美国,以便我们的人们可以使用它们并从中学习。
And we like stopping that would be such a restructuring of our Internet that it seems impossible.
而我们喜欢阻止那种情况,那将是对我们互联网的一次重大重组,似乎是不可能的。
Do you think maybe in that case, the big open weight models from China are actually a good thing in a sense,
你是否认为,也许在这种情况下,来自中国的大型开放权重模型实际上在某种意义上是件好事,
like for the U.S. companies because maybe the U.S. companies you mentioned earlier they are usually one generation behind in terms of what they release open source versus what they are using for example
比如对美国公司来说,因为也许你之前提到的美国公司,在开源发布的内容与例如他们正在使用的内容相比,通常落后一代,
one generation behind 常用搭配
落后一代
用于比较技术、产品或版本,表示比最新水平晚一代。
GPT OS as might not be the cutting edge model, Gemma 3 might not be, but they do that because they know this is safe to release,
GPT OS 可能不是最前沿的模型,Gemma 3 可能也不是,但他们这样做是因为他们知道发布这个是安全的,
but then when they see these companies see for example there is DeepSeek version 3.2 which is really awesome and it gets used and there is no backlash there is no security risk that could then again encourage them to release better models maybe that that in a sense is a very positive thing a hundred percent
但然后当他们看到这些公司,例如看到 DeepSeek 3.2 版本,它真的很棒,并且被使用,没有反弹,没有安全风险,这可能会再次鼓励他们发布更好的模型,也许那在某种意义上是件非常积极的事情,百分之百
these Chinese companies have set things into motion that I think would potentially not have happened if they were not all releasing models so
这些中国公司已经启动了某些事情,我认为如果他们没有全部发布模型,这些事情可能不会发生,所以
set things into motion 常用搭配
启动事情、使事情开始运转
指某个行动引发一连串后续发展。
I think that it's like I'm almost sure that those discussions have been had by leadership is there a possible future where the dominant models AI models
我认为这就像我几乎可以肯定那些讨论已经由领导层进行过,是否存在一个可能的未来,其中主导模型 AI 模型
in the world are all open source depends on the trajectory of progress that you predict.
世界上所有开源模型都取决于你预测的进步轨迹。
If you think saturation and progress is even coming within a few years, so essentially within the time where financial support is still very good,
如果你认为饱和和进步甚至会在几年内到来,那么基本上在资金支持仍然很好的时候,
then open models will be so optimized and so much cheaper to run that they will win out.
那么开源模型将被如此优化,运行成本如此之低,以至于它们会胜出。
win out 常用搭配
最终胜出
指经过竞争或比较后取得优势。
Essentially, this goes back to open source ideas where so many more people will be putting money into optimizing the serving of these open-weight common architectures that they will become standards.
本质上,这又回到了开源理念,即会有更多的人投入资金来优化这些开放权重通用架构的服务,使它们成为标准。
And then you could have chips dedicated to them and it'll be way cheaper than the offerings from these closed companies that are custom.
然后你可以有专用于它们的芯片,而且它会比这些定制闭源公司的产品便宜得多。
We should say that AI 27 report kind of predicts one of the things it does from a narrative perspective is that there will be a lot of centralization.
我们应该说,AI 27报告从叙事角度预测的一件事就是会有大量的集中化。
As the AI system gets smarter and smarter, the national security concerns will come to be and you'll centralize the labs and you'll become super secretive and there'll be this whole race from a military perspective of how to use between China and the United States.
随着AI系统变得越来越聪明,国家安全问题会出现,你会集中实验室,变得超级保密,并且从军事角度来看,会有这场关于如何在中美之间使用的全面竞赛。
And so all of this fun conversations we're having about LLMs,
所以,我们关于大语言模型的这些有趣对话,
the generals, the soldiers will come into the room and be like,
将军们、士兵们会走进房间,然后说:
all right, we're now in the Manhattan Project stage of this whole thing.
好吧,我们现在进入了这整件事的曼哈顿计划阶段。
I think 2025, 6, 7, 27, I don't think something like that is even remotely possible.
我认为2025年、6年、7年、27年,我觉得那样的事情根本不可能。
remotely possible 常用搭配
有一点点可能
常用于否定句,强调某事完全不可能。
I mean, you can make the same argument for computers, right?
我的意思是,你可以对电脑提出同样的论点,对吧?
make the same argument for 常用搭配
对……提出同样的论点
用于把之前的论证套用到另一个对象上。
You can say, okay, computers are capable and we don't want the general public to get them.
你可以说,好吧,电脑很有能力,我们不希望普通公众得到它们。
chips even ai chips but you see how like you know huawei makes chips now you know took a few years
芯片,甚至AI芯片,但你看到比如你知道华为现在制造芯片,你知道花了几年时间
but and i think that i don't think there is a way you can contain something like that like knowledge like that
但是,我认为,我不认为有办法遏制像那样的东西,像那样的知识
i think in this day and age it is impossible like the internet
我认为在当今这个时代,这是不可能的,就像互联网一样
in this day and age 常用搭配
在当今这个时代
用于谈论当前时代的情况,常带感慨或对比过去的语气。
i don't think this is a possibility on the manhattan project thing
我不认为在曼哈顿计划这件事上有这种可能性
one of my funny things making adam is i think that like a manhattan project like thing for open models would actually be pretty reasonable
我制作Adam时觉得有趣的一点是,我认为类似曼哈顿计划那样的开源模型项目实际上会相当合理
because wouldn't cost that much, but I think that that will come.
因为不会花费那么多,但我认为那会到来的。
It seems like culturally the companies are changing, but I agree with Sebastian on all the stuff that you just said.
看起来在文化上这些公司正在改变,但我同意Sebastian关于你刚才说的所有内容。
I agree with 句型
我同意某人的观点
I agree with [someone] on [something]
表达赞同某人看法时使用,后接人名或观点。
It's just like, I don't see it happening nor being helpful.
就像,我不认为这会发生,也不认为它有帮助。
I don't see it happening 句型
我不认为这会发生
I don't see [something] happening
口语中表达对某事不会发生的判断,后常接 nor/and 引出另一否定。
Yeah, I mean, the motivating force behind the Manhattan Project is there is civilizational risk.
是的,我的意思是,曼哈顿计划背后的驱动力是存在文明风险。
It's harder to motivate that for open source models.
对于开源模型来说,更难激发这种动力。
There's not civilizational risk.
不存在文明风险。
You think on the hardware side, we'll mention NVIDIA a bunch of times,
你认为在硬件方面,我们会多次提到英伟达,
Do you think Jensen and NVIDIA are going to keep winning?
你认为黄仁勋和英伟达会继续赢下去吗?
I think they have the downside that they have to iterate a lot and manufacture a lot.
我认为他们的劣势在于必须大量迭代和大量制造。
And I think they probably, what they're doing, they do innovate.
而且我认为他们可能,他们所做的,他们确实在创新。
But I think there's always the chance that there is something who does something fundamentally different, who gets very lucky and then does something.
但我认为总有可能出现某个做根本性不同事情的人,他非常幸运,然后做成了某事。
there's always the chance that 句型
总有可能……
there's always the chance that [clause]
用于表达某事虽然不确定但存在发生的可能性。
But the problem is, I think, adoption.
但问题在于,我认为,是采用。
You know, like the mode of NVIDIA is probably not just the GPU.
你知道,英伟达的模式可能不仅仅是GPU。
It's more like the CUDA ecosystem, and that has evolved over so many, I mean, two decades.
更像是CUDA生态系统,而且它已经演变了这么多,我是说,二十年。
I mean, even back when I was a grad student, I was in a lab where we did biophysical simulations, molecular dynamics, and we had a Tesla GPU back then just for the computation.
我是说,甚至在我读研究生的时候,我在一个实验室里做生物物理模拟、分子动力学,当时我们有一块Tesla GPU,就用来做计算。
It was 15 years ago now.
那是15年前的事了。
And they built this up for a long time, and that's the mode, I think.
他们花了很长时间建立这个,我认为这就是模式。
It's not the chip itself, although they have now the money to iterate and build and scale.
不是芯片本身,尽管他们现在有资金去迭代、构建和扩展。
But then it's really on the compatibility.
但接下来就真的取决于兼容性了。
It's like, well, if you're at that scale as a company,
就像,嗯,如果你作为一家公司达到了那种规模,
why would you go with something risky where it's only a few chips that they can make per year?
你为什么要选择有风险的东西,他们每年只能生产几块芯片?
why would you go with 句型
你为什么会选择……
why would you go with [something]
用反问表达某选择不合理,后接所选择的事物。
You go with a big one.
你会选择大厂。
But then I do think with LLMs now, also it will be easier to design something like CUDA, you know, like the next, so it took 15 years because it's hard.
但然后我确实认为,现在有了大语言模型,设计像CUDA这样的东西也会更容易,你知道,就像下一个,所以它花了15年,因为这很难。
But then now we have LLMs, we can maybe replicate CUDA.
但现在我们有了大语言模型,我们也许可以复制CUDA。
And I wonder if there will be a separation of the training and the inference compute,
我想知道训练和推理计算是否会分离,
as we kind of stabilize a bit more and more and more computers needed for inference.
随着我们逐渐稳定下来,推理所需的计算机越来越多。
That's supposed to be the point of the Grok acquisition.
这应该就是收购Grok的意义所在。
And and that's why part of what Vera Rubin is, where they have a new chip with no high bandwidth memory,
而这也是Vera Rubin的一部分原因,他们有一款没有高带宽内存的新芯片,
which is one of the, or very little, which is one of the most expensive pieces.
而高带宽内存是最昂贵的部件之一,或者他们只有很少。
It's designed for pre-fill,
它是为预填充设计的,
which is the part of inference where you essentially do a lot of matrix multiplications,
预填充是推理的一部分,你基本上要做大量的矩阵乘法,
and then you only need the memory when you're doing this autoregressive generation,
然后你只有在进行这种自回归生成时才需要内存,
and you have the KV cache swaps.
并且你有KV缓存交换。
So they have this new GPU that's designed for that specific use case,
所以他们有一款新的GPU,专为那个特定用例设计,
and then the cost of ownership per flop or whatever is actually way lower.
然后每flop或诸如此类的拥有成本实际上要低得多。
But I think that NVIDIA's fate lies in the diffusion of AI still.
但我认为NVIDIA的命运仍然在于AI的普及。
Their biggest clients are still these hyperscale companies,
他们最大的客户仍然是这些超大规模公司,
whether it's like Google, obviously can make TPUs.
比如谷歌,显然可以制造TPU。
Amazon is making Tranium.
亚马逊正在制造Tranium。
Microsoft will try to do its own things.
微软会尝试做自己的东西。
And like, so long as the pace of AI progress is high,
而且,只要AI进展的速度很高,
so long as 常用搭配
只要
引导条件从句,表示在某条件下某事成立。
NVIDIA's platform is the most flexible and people will want that.
NVIDIA的平台是最灵活的,人们会想要那个。
But if there's stagnation, then creating bespoke chips, there's more time to do it.
但如果出现停滞,那么创建定制芯片就有更多时间去做。
It's interesting that NVIDIA is quite active in trying to develop all kinds of different products.
有趣的是,NVIDIA非常积极地尝试开发各种不同的产品。
They tried to create areas of commercial value that will use a lot of GPUs.
他们试图创造能够使用大量GPU的商业价值领域。
But they keep innovating and they're doing a lot of incredible research.
但他们不断创新,并且正在进行大量令人难以置信的研究。
Everyone says the company is super oriented around Jensen and how operationally plugged in he is.
每个人都说这家公司非常以Jensen为中心,以及他在运营上多么深入参与。
And it sounds so unlike many other big companies that I've heard about.
这听起来和我听说过的许多其他大公司非常不同。
And so long as that's the culture, I think that I will expect them to keep progress happening.
只要那是文化,我想我会期望他们继续取得进展。
And it's like he's still in the Steve Jobs era of Apple.
这就像他仍然处于苹果的史蒂夫·乔布斯时代。
So long as that is how it operates, I'm pretty optimistic for their situation.
只要它这样运作,我对他们的情况相当乐观。
Because it's like it is their top order problem.
因为这就好像是他们最优先的问题。
And I don't know if making these chips for the whole ecosystem is the top goal of all these other companies.
而且我不知道为整个生态系统制造这些芯片是否是所有其他公司的首要目标。
They'll do a good job, but it might not be as good of a job.
他们会做得不错,但可能不会那么好。
Since you mentioned Jensen, I've been reading a lot about history and about singular figures in history.
既然你提到了Jensen,我一直在读很多关于历史和历史上杰出人物的内容。
What do you guys think about the single man, woman view of history?
你们怎么看历史中的单一男人、女人视角?
How important are individuals for steering the direction of history in the tech sector?
个人在引导科技领域历史方向上有多重要?
So, you know, what's NVIDIA without Jensen?
所以,你知道,没有Jensen的NVIDIA是什么?
You mentioned Steve Jobs. What's Apple without Steve Jobs?
你提到了Steve Jobs。没有Steve Jobs的Apple是什么?
What's XAI without Elon?
没有Elon的XAI是什么?
Or DeepMind without Demis?
或者没有Demis的DeepMind?
People make things earlier and faster, where scientifically, many great scientists credit to being in the right place at the right time and still making the innovation,
人们更早、更快地创造事物,而在科学上,许多伟大的科学家将其归功于在正确的时间处于正确的位置并仍然进行创新,
being in the right place at the right time 常用搭配
在正确的时间处于正确的位置(即运气好)
形容成功部分归因于时机和运气,常用于谈论机遇。
where eventually someone else will still have the idea.
而最终其他人仍然会有这个想法。
So I think that in that way, Jensen is helping manifest this GPU revolution much faster and much more focused than without having a person there it would do
所以我认为,通过这种方式,Jensen正在帮助更快、更集中地实现这场GPU革命,比没有人在那里推动要快得多、专注得多。
and this is making the whole AI build out faster
而这正在使整个人工智能建设更快。
but I do still think that eventually like something like ChatGPT would have happened in a build out like this
但我仍然认为,最终像ChatGPT这样的东西会在这样的建设中出现。
would have happened but it probably would not have been as fast or like as like I think that's the sort of flavor that is applied
它会发生,但可能不会那么快,或者像我认为的那样,这就是所应用的那种风格。
people these individual people are people who are placing bets on something some get lucky some don't
这些人,这些个体,是在对某件事下注的人,有些人幸运,有些人不。
but if you don't have these people at the helm it will be more diffused
但如果你没有这些人掌舵,它会更分散。
at the helm 常用搭配
掌舵,处于领导地位
比喻某人负责领导或管理某个组织或项目。
it's almost like investing in a ETF versus individual stocks individual stocks might go up might go down more heavily than an ETF which is more balanced
这几乎就像投资ETF与个股,个股可能比更平衡的ETF涨跌更剧烈。
it will eventually go up over time we'll get there but it's just like you know
它最终会随着时间的推移而上涨,我们会达到那里,但就像你知道的。
like focus I think is the thing passion and focus
就像专注,我认为是关键,激情和专注。
isn't there a real case to be made that without Jensen there's not a reinvigoration of the deep learning revolution
难道没有一个真实的理由表明,没有Jensen,就不会有深度学习革命的复兴吗?
it could have been 20 years later is the thing that I would say yeah yeah to
它可能会晚20年,这就是我会说“是的,是的”的事情。
a 20 or like another AI when like a deep learning winter
一个20年或像另一次AI寒冬,当像深度学习寒冬
could have come. Yeah, if GPUs weren't around,
可能会到来。是的,如果GPU不存在,
that could change history completely,
那可能会完全改变历史,
because you could think of all the other technologies
因为你可以想到所有其他技术
that that could have come in the meantime,
那可能在此期间出现,
and the focus of human civilization could the silicon value would be captured by different hype,
而人类文明的焦点可能,硅谷的价值会被不同的炒作所捕获,
but I do think it is, I mean,
但我确实认为它是,我的意思是,
there's certainly an aspect where it was all planned,
肯定有一个方面是这一切都是计划好的,
the GPU trajectory, but on the other hand,
GPU的发展轨迹,但另一方面,
on the other hand 常用搭配
另一方面
用于引出与前面观点相对或补充的另一个角度,口语和书面都常用。
it's also a lot of lucky coincidences, for example,
也有很多幸运的巧合,例如,
for example 常用搭配
例如
用于举例说明前面的观点,口语中常插入句中使用。
or good intuition, like the investment into this, let's say, biophysical simulations,
或者好的直觉,比如对这方面的投资,比如说,生物物理模拟,
let's say 地道口语
比如说
口语中用来举例或提出一个假设的例子,语气随意。
or like, I mean, I think it started with video games,
或者像,我的意思是,我认为它始于视频游戏,
and then it just happened to be good at linear algebra,
然后它恰好擅长线性代数,
happened to be 常用搭配
恰好是,碰巧是
表示某事是偶然发生的,而非计划之中。
because video games require a lot of linear algebra,
因为视频游戏需要大量的线性代数,
and then you have the biophysical simulations,
然后你有生物物理模拟,
and then but still I don't think the plan, the master plan, was AI, I think
然后但仍然,我不认为计划,那个总体规划,是AI,我认为
there was just it happened to be Alex Kraszewski,
只是恰好是Alex Kraszewski,
so someone took these GPUs and like, hey, let's try to train a neural network on that,
所以有人拿了这些GPU,然后像,嘿,让我们试着在那上面训练一个神经网络,
and happen to work really well, and I think it only happened
并且恰好效果非常好,我认为它之所以发生
happen to work really well 常用搭配
恰好效果非常好
用 happen to 表示某事碰巧发生或结果很好,常与动词连用。
because you could purchase those GPUs, gaming would have
是因为你可以购买那些GPU,游戏本来会
created a demand for faster processors.
创造了对更快处理器的需求。
If Nvidia had gone out of business in the early days,
如果英伟达在早期就倒闭了,
gone out of business 常用搭配
倒闭,停业
指公司因经营不善而停止运营,常用于谈论企业命运。
that's what I would think like,
那我会觉得,
I think that the GPUs would have been different for the Alex,
我认为GPU对Alex来说会有所不同,
but I think like GPUs would still exist at the time of AlexNet
但我认为像GPU在AlexNet时代仍然会存在,
and at the time of the transformer.
在Transformer时代也是如此。
It was just hard to know
只是很难知道
if it would be one company as successful or multiple smaller companies with worse chips.
它是否会成为一家同样成功的公司,还是多家拥有更差芯片的小公司。
But I don't think that's like a 100 year delay.
但我不认为那是100年的延迟。
It might be a decade delay.
可能是十年的延迟。
Well, it could be one, two, three, four, five decade delay.
嗯,可能是1、2、3、4、5十年的延迟。
I mean, I just can't see Intel or AMD doing what Nvidia did.
我的意思是,我就是无法想象英特尔或AMD会做英伟达所做的事。
I just can't see 句型
我就是无法想象/看不出
I just can't see [someone/something] doing [something]
用于表达强烈认为某事不可能或不会发生,后接名词或动名词。
I don't think it would be a company that exists.
我不认为它会是一家存在的公司。
I think it would be a different company would rise like Silicon Graphics or something.
我认为会有一家不同的公司崛起,像硅谷图形之类的。
So yeah, some company that has died would have done it.
所以是的,某个已经消亡的公司可能会做到。
But it does like just look looking at it.
但它确实就像只是看着它。
It seems like these singular figures, these leaders have a huge impact on the trajectory of the world.
看起来这些非凡的人物,这些领导者对世界的轨迹有着巨大的影响。
Obviously incredible teams behind them,
显然他们背后有不可思议的团队,
but you know, having that kind of very singular, almost dogmatic focus is necessary to make progress.
但你知道,拥有那种非常独特、几乎教条式的专注对于取得进展是必要的。
Yeah, I mean, even with
是的,我的意思是,即使有
a GPT. It wouldn't exist if there wasn't a person, Ilya, who pushed for this scaling, right?
一个GPT。如果不是有一个人,伊利亚,推动了这种规模化,它就不会存在,对吧?
I mean, yeah, Dario is also deeply involved in that.
我的意思是,是的,达里奥也深度参与了那件事。
You read some of the histories of OpenAI.
你读过一些OpenAI的历史。
It almost seems wild thinking about how early these people were like,
想想这些人当时有多早就提出,几乎显得很疯狂,比如,
we need to hook up 10,000 GPUs and take all of OpenAI's compute and train one model.
我们需要连接1万个GPU,拿走OpenAI所有的算力,训练一个模型。
There's a lot of people there that didn't want to do that.
那里有很多人不想那样做。
Which is an insane thing to believe.
相信这一点是件疯狂的事。
To believe scaling before scaling has any indication that is going to materialize.
在规模化有任何迹象表明它会实现之前就相信规模化。
Again, singular figures.
再次强调,独特的人物。
Speaking of which, 100 years from now, this is presumably post-singularity, whatever singularity is,
说到这个,从现在起100年后,这大概是后奇点时代,不管奇点是什么,
Speaking of which 地道口语
说到这个,顺便提一下
口语中用来引出与刚提到的话题相关的新内容。
when historians look back at our time now,
当历史学家回顾我们现在这个时代,
what technological breakthroughs would they really emphasize as the breakthroughs that led to the singularity?
他们真正会强调哪些技术突破是导致奇点的突破?
So, so far we have Turing to today, 80 years.
所以,到目前为止,我们从图灵到今天,80年。
I think it would still be computing. like the umbrella term computing,
我认为它仍然会是计算。就像计算这个总称,
just, I don't necessarily think it's even, like, 100 years, 200 years from now, it would be AI.
只是,我不一定认为甚至像100年、200年后,它会是AI。
It could still well be computers, you know.
它很可能仍然是计算机,你知道。
Just, we are now taking better advantage of computers, but, like, the fact of computing.
只是,我们现在更好地利用了计算机,但是,就像,计算这个事实。
It's basically Moore's Law kind of discussion.
这基本上是摩尔定律之类的讨论。
You're not, even the details of CUDA and GPUs won't even be remembered.
你不会,甚至CUDA和GPU的细节都不会被记住。
And it won't be all this software turmoil.
而且不会全是这种软件上的混乱。
It'll be just, obviously, compute.
显然,它只会是计算。
I generally agree, but it's like, is the connectivity of the internet and compute able to be merged?
我大体上同意,但就像,互联网的连通性和计算能力能融合在一起吗?
Or is it both of them?
还是说两者都是?
I think the internet will probably be related to, yeah, I mean, communication.
我觉得互联网可能跟,对,我是说,通信有关。
It could be a phone, internet, satellite, that stuff.
它可能是电话、互联网、卫星,诸如此类的东西。
Where, yeah, and compute is more like the scaling aspect of it.
而,对,计算更像是它的扩展方面。
It's possible that the internet is completely forgotten.
有可能互联网会被完全遗忘。
The internet is wrapped into the phone networks, like communication networks.
互联网被融入到电话网络中,就像通信网络一样。
This is just another manifestation of that.
这只是那种情况的另一种体现。
And the real breakthrough comes from just the increased compute, is the Moore's Law, broadly defined.
而真正的突破来自于计算能力的提升,也就是广义上的摩尔定律。
Well, I think that connection of people is very fundamental to it.
嗯,我觉得人与人之间的连接对此非常根本。
So it's like, you can talk to anyone, you want to find the best person in the world or something, they are somewhere in the world.
所以就像,你可以和任何人交谈,你想找到世界上最好的人之类的,他们就在世界上的某个地方。
And being able to have that flow of information, the AIs will also rely on this.
而能够拥有那种信息流动,AI 也会依赖这一点。
I think I've been fixating on the like, when I said the dream was dead about the one central model.
我觉得我一直纠结于,就像,当我说那个梦想已死,关于那个单一中心模型。
And the thing that is evolving is like people have many agents for different tasks.
而正在演变的是,人们为不同的任务拥有许多智能体。
People always start doing this with different clods for different tasks.
人们总是开始用不同的工具来做不同的任务。
And it's described as many AGIs in the data center where each one manages and they talk to each other.
它被描述为数据中心里的许多 AGI,每一个都负责管理,并且它们互相交流。
And like that is so reliant on networking and free flow of information on top of compute.
就像那样,它非常依赖网络和计算之上的信息自由流动。
But like networking, especially with GPUs, is such a part of scaling up compute.
但就像网络,尤其是用GPU时,是扩展计算的重要组成部分。
Like the GPUs and the data centers need to talk to each other.
就像GPU和数据中心需要互相通信。
Anything about neural networks will be remembered?
关于神经网络的一切会被记住吗?
Like, do you think there's something very specific and singular to the fact that it's neural networks that's seen as a breakthrough,
比如,你觉得神经网络被视为突破这件事,有什么非常具体而独特的地方吗?
like a genius that you're basically replicating in a very crude way the human mind, the structure of the human brain, the human mind?
就像一个天才,你基本上是在以一种非常粗糙的方式复制人类心智、人脑的结构、人类心智?
I think without the human mind, we probably wouldn't have neural networks.
我认为没有人类心智,我们可能就不会有神经网络。
because it just was an inspiration for that.
因为它只是那方面的灵感来源。
But at the other end, I think it's just so different.
但在另一端,我觉得它就是如此不同。
at the other end 常用搭配
在另一方面
用于对比讨论中的另一端或另一面,口语中常见。
I mean, it's digital versus, you know, biological,
我是说,它是数字的,而那是,你知道,生物的,
I mean 地道口语
我是说
口语中用来解释、补充或修正自己刚说的话。
that I do think it will probably be more like grouped as an algorithm.
我确实认为它可能更像被归类为一种算法。
That's massively paralyzable on this particular kind of compute.
那在这种特定类型的计算上是高度可并行化的。
Could have well been like genetic computing, like genetic algorithms just as paralyzed a thing.
本来也可能是像基因计算,像遗传算法一样可并行化的东西。
It just happens that this is more efficient, works better, you know.
只是恰好这个更高效、效果更好,你知道。
It just happens that 句型
只是恰好……
It just happens that [clause]
用于说明某事是偶然发生的,而非刻意安排。
And it very well could be that the LLM, you know, the neural networks, the way we architect them now,
而且很可能,LLM,你知道,神经网络,我们现在构建它们的方式,
It's just a small component of the system that leads to singularity.
它只是通向奇点的系统中的一个小组件。
The thing is, if you think of it 100 years,
问题是,如果你想想100年后,
The thing is 地道口语
问题是;关键在于
口语中引出重点或困难所在,常用于解释情况。
society, I think, can be changed more with more compute intelligence
我认为,社会可以通过更多的计算智能发生更大的改变,
because of autonomy.
因为有了自主性。
But looking at this, what are the things from the Industrial Revolution that we remember?
但看看这个,我们记得的工业革命中的东西是什么?
We remember the engine is probably the equivalent of the computer in this.
我们记得发动机大概就相当于这里的计算机。
But there's a lot of other physical transformations that people are aware of,
但还有很多其他人们所知道的物理变革,
like all the cotton gin and all these things that these machines
比如所有的轧棉机,以及所有这些机器,
that are still known air conditioning refrigerators
仍然为人所知的空调、冰箱,
like some of these things from ai will still be known
就像人工智能中的一些东西仍然会为人所知,
like the word transformer could still very well be known i would guess
比如“transformer”这个词很可能仍然会为人所知,我猜,
that deep learning is definitely still known
深度学习肯定仍然会为人所知,
but the transformer might be evolved away from in 100 years of with asiai researchers everywhere.
但transformer可能会在100年后被淘汰,随着各地都有亚洲人工智能研究人员。
But I think deep learning is likely to be a term that is remembered.
但我认为深度学习很可能是一个会被记住的术语。
And I wonder what the air conditioning and the refrigeration of the future is that AI brings.
我想知道人工智能带来的未来的空调和制冷是什么。
Is there, if we travel forward 100 years from now, we transport there right now,
有没有,如果我们从现在向前穿越100年,我们现在就传送到那里,
what do you think is different? How do you think the world looks different?
你觉得有什么不同?你觉得世界看起来有什么不同?
First of all, you think there's humans? You think there's robots everywhere walking around?
首先,你觉得有人类吗?你觉得到处都是机器人在走动吗?
I do think specialized robots for sure for certain tasks.
我确实认为对于某些任务肯定会有专门的机器人。
Humanoid form?
人形形态?
That I'm maybe half humanoid. We'll see.
我可能算半个类人机器人吧。我们走着瞧。
We'll see 地道口语
我们走着瞧;到时候看
口语中表示对未来的事不确定,等以后再看结果。
I think for certain things, yes, there will be humanoid robots because it's just amenable for the environment.
我觉得在某些事情上,是的,会有类人机器人,因为它就是很适合这个环境。
But for certain tasks, it might make sense.
但对于某些任务,它可能是有道理的。
make sense 常用搭配
有道理;说得通
用于表示某个做法或想法合理、可行。
What's harder to imagine is how we interact with devices and what humans do with devices.
更难想象的是我们如何与设备互动,以及人类用设备做什么。
I mean, I'm pretty sure it will probably not be the cell phone, will probably not be the laptop, will it be implants.
我是说,我很确定它很可能不会是手机,很可能不会是笔记本电脑,会不会是植入体呢。
I mean, it has to be brain, computer, and devices, right?
我是说,它必须是大脑、计算机和设备,对吧?
I mean, 100 years from now, it has to, like, given the progress we're seeing now, there has to be, unless there's legitimately...
我是说,从现在起100年后,它必须,就像,鉴于我们现在看到的进步,必须有,除非真的有……
complete alteration of how we interact with reality.
彻底改变我们与现实互动的方式。
On the other hand, if you think of cars, cars are older than 100 years, right?
另一方面,如果你想想汽车,汽车已经超过100年了,对吧?
On the other hand 常用搭配
另一方面
用于引出与前面观点相对或补充的另一方面。
And it's still the same interface.
而且它仍然是同样的界面。
We haven't replaced cars with something else.
我们没有用别的东西取代汽车。
We just made the cars better, but it's still steering wheel, it's still wheels, you know?
我们只是把汽车做得更好了,但它仍然是方向盘,仍然是轮子,你知道吗?
I think we'll still carry around a physical brick of compute because people want some ability to have a private,
我觉得我们仍然会随身携带一块实体的计算砖块,因为人们想要某种能力来拥有一个私密的,
like you might not engage with it as much as a phone, but having something where you could have private information that is yours as an interface
就像你可能不会像用手机那样频繁地使用它,但拥有一个你可以拥有属于你自己的私人信息作为界面的东西
between the rest of the internet, I think is something that people will still exist.
在互联网的其他部分之间,我认为这是人们仍然会存在的东西。
It might not look like an iPhone and it might be used a lot less, but I still expect to have people carry things around.
它可能看起来不像iPhone,可能用得少得多,但我仍然期望人们会随身携带东西。
Why do you think the smartphone is the embodiment of private?
为什么你认为智能手机是私密的化身?
There's a camera on it.
上面有摄像头。
Private for you, like encrypted messages, encrypted photos, you know what your life is.
对你来说是私密的,比如加密信息、加密照片,你知道你的生活是什么样。
you know what your life is 地道口语
你知道你的生活是什么样的
口语中用来强调对方明白自己所指的私人生活内容。
I guess this is a question on how optimistic on brain-machine interfaces you are.
我想这是一个关于你对脑机接口有多乐观的问题。
is all that just going to be stored in the cloud and your whole calendar like it it's hard to think about processing
所有这些都只是存储在云端,你的整个日历,就像很难想象处理
all the information that we can process visually through brain machine interfaces presenting something like a calendar or something to you like it's hard to just think about
所有我们能通过脑机接口视觉处理的信息,比如向你呈现日历或类似的东西,就像很难只是想象
knowing without looking you know your email inbox like you signal to a computer and then you just know your email inbox like what
不用看就知道,你知道你的电子邮件收件箱,就像你向电脑发信号,然后你就知道你的电子邮件收件箱,像什么
does that like is that something that the human brain can handle being piped into it non-visually.
那像那样,是大脑能处理非视觉输入的东西吗?
I don't know exactly how those transformations happen.
我不确切知道那些转换是如何发生的。
Because humans aren't changing in 100 years.
因为人类在100年内不会改变。
I think agency and community are things that people actually want.
我认为能动性和社区是人们真正想要的。
Local community, yeah. So people you are close to, being able to do things with them and being able to describe meaning to your life and to be able to do things.
本地社区,是的。所以你亲近的人,能够和他们一起做事,能够描述你生活的意义,并且能够做事。
I think that that is... If not in 100 years, I don't think that human biology is changing away from those on a timescale that we can discuss.
我认为那是……如果不是在100年内,我不认为人类生物学会在我们能讨论的时间尺度上偏离那些。
And I think that UBI does not solve agency.
而且我认为全民基本收入不能解决能动性。
I do expect mass wealth,
我确实期待大众财富,
and I hope that it is spread so that the average life does look very different in 100 years.
我希望它能被广泛分享,这样普通人的生活在一百年后看起来会非常不同。
But that's still a lot to happen in 100 years.
但在一百年里仍然会发生很多事。
If you think about countries that are early in their development process to getting access to computing and internet,
如果你想想那些在发展进程早期就获得计算和互联网接入的国家,
like to build all the infrastructure and to have policy that shares one nation's wealth with another is, I think it's an optimistic view to see all of that happening in 100 years while they are still independent entities and not just like absorbed into some international order by force.
喜欢建设所有基础设施,并制定政策将一个国家的财富与另一个国家分享,我认为,看到这一切在一百年内发生,同时它们仍然是独立实体,而不是像被武力吸收进某种国际秩序,这是一种乐观的看法。
But there could be just better, more elaborate, more effective social support systems that help alleviate some levels of basic suffering from the world.
但可能会有更好、更精细、更有效的社会支持系统,帮助减轻世界上一些基本层面的痛苦。
You know, the transformation of society where a lot of jobs are lost in the short term.
你知道,社会转型中很多工作在短期内会消失。
I think we have to really remember that each individual job that's lost is a human being who's suffering.
我认为我们必须真正记住,每一个失去的工作都是一个正在受苦的人。
That's like when jobs are lost, it scales a real tragedy.
就像当工作失去时,它演变成一场真正的悲剧。
You can make all kinds of arguments about economics or it's all going to be okay. It's good for the GDP.
你可以提出各种关于经济学的论点,或者一切都会好起来的。这对GDP有好处。
There's going to be new jobs created.
会有新的工作被创造出来。
Fundamentally, the individual level for that human being, that's real suffering.
从根本上说,对那个人来说,个体层面,那是真正的痛苦。
That's a real personal sort of tragedy, and we have to not forget that as the technologies are being developed.
那是一种真正的个人悲剧,随着技术的发展,我们不能忘记这一点。
And also my hope for all the AI slop we're seeing is that
而且我对我们看到的这些AI垃圾内容的希望是
there will be a greater and greater premium for the fundamental aspects of the human experience
人类体验的基本方面会获得越来越高的溢价
that are like in person, the things that we all, like seeing each other, talking together in person.
比如面对面的,我们所有人都有的,比如见面、面对面交谈。
The next few years are definitely going to be an increased value on physical goods and events
未来几年,实物商品和活动肯定会越来越有价值
and even more pressure on slop.
而垃圾内容会面临更大的压力。
So the slop is only starting.
所以垃圾内容才刚刚开始。
The next few years will be more and more diverse. versions of slop they would be drowning in slop
未来几年会有越来越多样化的垃圾内容版本,他们会淹没在垃圾内容中
so i'm hoping that we society drowns in slop enough to snap out of it and be like we can't like none like it just doesn't matter we all can't deal with it and
所以我希望我们社会淹没在垃圾内容中足够多,从而摆脱它,然后说我们不能像没有一样,它只是不重要,我们所有人都无法应对它,然后
snap out of it 常用搭配
摆脱(某种状态);振作起来
口语中表示从消极、沉迷或不正常的状态中恢复过来。
then like the physical has such a higher premium on it even like uh classic examples
然后就像实物会有更高的溢价,甚至像呃经典的例子
higher premium on it 常用搭配
对它收取更高的溢价
谈论某物因稀缺或特殊而比同类更值钱时使用。
i honestly think this is true and i think we'll get tired of it we are already kind of tired of it same
我真心认为这是真的,而且我觉得我们会厌倦它,我们已经有点厌倦它了,同样
get tired of it 常用搭配
对它感到厌倦
表示对某事物逐渐失去兴趣或耐心,日常口语常用。
with i mean even art i don't think art will go away i mean you have paintings physical paintings there's more value
同样,我是说甚至艺术,我不认为艺术会消失,我是说你有画作,实体的画作,有更多价值
go away 常用搭配
消失、不复存在
谈论某事物是否会继续存在时使用,口语常用。
not just monetary value but just more value appreciation for something that is the actual painting than a photocopy of that painting
不仅仅是金钱价值,而是对实际画作比那幅画的复印件有更多的价值欣赏
not just 句型
不仅仅是
not just [X] but [Y]
用于强调不止一个方面,后面常接 but also。
it could be a perfect digital reprint of that but there is something when you go to a museum and you look at that art and you see that real thing and you think about okay a human i
它可能是完美的数字重印,但当你去博物馆,看着那件艺术品,看到真实的东西,然后你想,好吧,一个人类,我
think about 常用搭配
思考、琢磨
表示停下来考虑某事,日常口语常用。
don't know it's like a craft you have like appreciation for that
不知道,这就像一门手艺,你对此有某种欣赏
and I think the same is true for writing, for talking, for any type of experience
我认为写作、谈话、任何类型的体验也是如此
the same is true for 句型
对……也是如此
the same is true for [something]
用于把前面说的情况类推到其他对象上。
where it will be, I do unfortunately
它会变成,我确实不幸地认为
think it will be like a dichotomy
认为它会像一种二分法
like it will be like a fork where well
就像它会像一个分叉,嗯
some things will be automated like you know there are not as many paintings as they used to be 200 years ago
有些东西会被自动化,比如你知道,绘画不像200年前那么多了
used to be 句型
过去曾经是
used to be [something]
表示过去存在但现在已经不存在的状态或情况。
there are more, more photographs, more photocopies
有更多,更多照片,更多复印件
but at the same time it won't go away, there
但与此同时它不会消失,那里
at the same time 常用搭配
与此同时
用于补充与前一句形成对照或并列的信息。
will be a, you know, value in that
会有一种,你知道,价值在其中
I think that the difference will just be a bit, you know, what's the proportion of that
我认为区别只是有点,你知道,那比例是多少
but personally I, I have a hard time reading things
但就个人而言,我,我很难阅读东西
have a hard time 常用搭配
做某事很困难、很费劲
后接动名词,表示做某事感到吃力或不情愿。
where I obviously see it's, um, obviously AI generated
当我明显看到它是,嗯,明显是AI生成的时候
I'm like sorry that might be really good information there but I have like a certain nah not for me I think
我会觉得抱歉,那可能是非常好的信息,但我有一种,不,不适合我,我认为
not for me 地道口语
不适合我、我不感兴趣
口语中委婉表达自己不喜欢或不接受某事物。
eventually they'll fool you and it'll be on platforms
最终它们会骗过你,而且它会出现在平台上
that give ways of verifying or building trust
这些平台提供验证或建立信任的方式
so you will trust that Lex is not AI generated, having been here
所以你会相信Lex不是AI生成的,因为他一直在这里
so then you have trust in this channel
那么你就信任这个频道
have trust in 常用搭配
信任……
表示对某人或某机构有信心、信赖。
but it's harder for new people that don't have that trust
但对没有那种信任的新人来说更难
well that will get interesting because I think fundamentally I think there's a solvable problem
嗯,那会变得有趣,因为我认为从根本上说,我认为有一个可解决的问题
by having, you know, trust in certain outlets that they won't do it but
通过,你知道,信任某些渠道,它们不会这样做,但
It's all going to be kind of trust-based there.
这将会是一种基于信任的模式。
There will be some systems to authorize.
会有一些系统来进行授权。
Okay, this is real. This is not real.
好,这是真的。这不是真的。
There will be some tell... tell... tell... science...
会有一些告诉……告诉……告诉……科学……
where you can obviously tell this is AI-generated and this is not, but
你可以明显看出这是AI生成的,而这不是,但是
tell this is 句型
分辨出这是……
can tell [this] is [something]
用于表示能辨别出某事物的性质,常与 can 连用。
they want, I mean, some will be so good that it's hard to tell, and
他们想要,我的意思是,有些会好到很难分辨,然后
hard to tell 常用搭配
很难分辨
表示难以区分或判断某事物,口语常用。
then you have to trust, and, um, well, that... that will get interesting and a bit problematic.
然后你就必须信任,而且,嗯,好吧,那……那会变得有趣,也有点问题。
The extreme case of this is to watermark all human content, so
极端的情况是给所有人类内容打上水印,所以
all photos that we take on our own have some watermark until they are edited or something like
我们自己拍的所有照片都有某种水印,直到它们被编辑或类似处理
this and software can manage communications with the device manufacturer to maintain like human editing, which is the opposite of the discussion to
这样,软件可以与设备制造商沟通,以维持类似人类编辑,这与讨论的相反,即
try to watermark AI images, and then you can make a Google image that has a watermark and use a different Google tool to remove the water...
试图给AI图像打水印,然后你可以制作一个带有水印的谷歌图像,并使用不同的谷歌工具去除水印……
Yeah, yeah, it's going to be mom's race. Yeah, uh, and we've been mostly focusing on the positive aspects of AI.
是的,是的,这将是妈妈的竞赛。是的,呃,我们一直主要关注AI的积极方面。
I mean, there's also the, all the capabilities we've been talking about can be used to destabilize human civilization
我的意思是,还有,我们一直在谈论的所有能力都可以被用来破坏人类文明的稳定
with even just relatively dumb AI applied at scale.
即使只是相对愚蠢的AI大规模应用。
at scale 常用搭配
大规模地
表示以很大的规模进行某事,常用于商业或技术语境。
And then further and further, super intelligent AI systems.
然后越来越进一步,超级智能AI系统。
Of course, there's the sort of do-mer take that's important to consider a little bit as we develop these technologies.
当然,有一种末日论者的观点,在我们开发这些技术时,稍微考虑一下这一点很重要。
What gives you hope about the future of human civilization?
对于人类文明的未来,什么让你抱有希望?
Everything we've been talking about, are we going to be okay?
我们一直在谈论的这一切,我们会没事吗?
I think we, we will.
我觉得我们,我们会的。
I'm, I'm definitely a worrier, both about AI and non-AI things.
我,我绝对是个爱担心的人,既担心AI的事,也担心非AI的事。
But, um, humans do tend to find a way.
但是,嗯,人类确实往往会找到办法。
find a way 常用搭配
找到办法
表示在困难中设法解决问题,口语常用。
I think that's what humans are built for is to have community and find a way to figure out problems.
我觉得这正是人类的天性所在,就是拥有社群,并找到解决问题的办法。
figure out 常用搭配
弄明白、解决
表示通过思考或尝试找到答案或解决办法。
And that's what has gotten us to this point.
而正是这一点让我们走到了今天。
gotten us to this point 常用搭配
让我们走到今天这一步
回顾过去,说明某因素促成了当前局面。
And I think that the AI opportunity in related technologies is really big.
而且我认为,AI及相关技术带来的机遇真的很大。
And I think that there's big social and political problems to help everybody understand.
而且我认为,存在着重大的社会和政治问题,需要帮助所有人理解。
That and I think that that's what we're staring at a lot of right now.
这一点,而且我觉得这正是我们现在正面对着的许多问题。
staring at 常用搭配
面对着、正视着
表示正面临某个问题或局面,带有不得不面对之意。
It's like the world is a scary place and AI is a very uncertain thing.
就好像这个世界是个可怕的地方,而AI是一件非常不确定的事。
And it takes a lot of work that is not.
而这需要大量的工作,而这并不是。
Necessarily building things, it's like telling people and understanding people that the people building AI are historically not motivated or wanting to do that.
未必是建造东西,而是像告诉人们并让人们理解,那些建造AI的人历来并没有动力或意愿去做这件事。
It is something that is probably doable.
这件事大概是可行的。
Just will take longer than people want.
只是会比人们希望的要花更长时间。
We have to go through that long period of like hard, distraught AI discussions.
我们必须经历那段漫长的、充满艰难和焦虑的AI讨论期。
If we want to have the lasting benefits.
如果我们想要获得持久的益处。
Yeah, through that process.
是的,通过那个过程。
I'm especially excited that we get a chance, uh, to better understand ourselves.
我特别兴奋的是,我们有机会,呃,更好地了解我们自己。
Also at the individual level as humans and at the civilization level.
既在作为人类的个体层面上,也在文明的层面上。
It answers some of the big mysteries, like.
它解答了一些重大的谜团,比如。
What is this whole like consciousness thing going on here?
这里发生的整个意识之类的事情到底是什么?
Seems to be truly special, like there's a real miracle in our mind.
似乎真的很特别,就像我们头脑中有一个真正的奇迹。
And AI puts a mirror to ourselves and get to answer some of the big questions about
而AI为我们提供了一面镜子,让我们能够回答一些关于……的大问题
puts a mirror to 常用搭配
为……提供一面镜子、让人反思
比喻某事物促使人反观自身、认识自己。
like what, what is this whole thing going on here?
比如,这里发生的整个事情到底是什么?
Well, one thing about that is also what I do think makes us very different from AI,
嗯,关于这一点,还有一点就是我认为是什么让我们与AI非常不同,
and why I don't worry about AI taking over is, like you said,
以及为什么我不担心AI接管,就像你说的,
consciousness, we humans, we decide what we want to do. AI in its current implementation,
意识,我们人类,我们决定我们想做什么。AI在当前的实现中,
I can't see it changing. You have to tell it what to do,
我看不到它会改变。你必须告诉它该做什么,
and so you have still the agency. It doesn't take the agency from you because
所以你还拥有自主权。它不会夺走你的自主权,因为
you have to, you just, it becomes a tool you can think of it as a tool.
你必须,你只是,它变成了一个工具,你可以把它看作一个工具。
think of it as 句型
把它看作……
think of [something] as [something]
用于把某事物重新定义或归类为另一事物。
You tell it what to do, it will be more automatic than other previous tools.
你告诉它该做什么,它会比之前的其他工具更自动化。
It's certainly more powerful than a hammer.
它肯定比锤子更强大。
It can figure things out, but it's still you in in in charge, right?
它能解决问题,但仍然是你在负责,对吧?
in charge 常用搭配
负责、掌管
表示对某事有控制权或责任,常与 be 连用。
So the AI is not in charge. You're in charge.
所以AI不是负责人。你才是负责人。
You tell the AI what to do, and it's doing it for you.
你告诉AI该做什么,它就在为你做。
So in the post-singularity, post-apocalyptic war between humans and machines, you're saying humans are worth fighting for?
所以在后奇点、后末日的机器与人类战争中,你是说人类值得为之战斗?
worth fighting for 常用搭配
值得为之奋斗
表示某事物重要到值得付出努力去争取或保护。
100%. I mean, this is the movie Terminator they made in the 80s, essentially.
100%。我的意思是,这基本上就是他们在80年代拍的电影《终结者》。
And I do think, well, the only thing I can see going wrong is, of course,
我确实认为,嗯,我能看到的唯一出问题的地方,当然是,
going wrong 常用搭配
出问题、出错
表示事情发展不顺利或出现故障。
if things are explicitly programmed to do the thing that is harmful, basically.
如果事物被明确编程去做有害的事情,基本上就是这样。
I think actually in that Terminator type of setup, I think humans win.
我认为实际上在那种终结者式的设定中,我认为人类会赢。
Mm-hmm. I think we're too clever.
嗯。我认为我们太聪明了。
It's hard to explain how we figure it out, but we do.
很难解释我们是如何弄明白的,但我们确实能做到。
figure it out 常用搭配
弄明白、找到解决办法
表示暂时不知道答案或做法,但会通过思考或尝试解决,口语常用。
And we'll probably be using local LLMs, open-source LLMs, to help fight the machines.
我们可能会使用本地的大语言模型、开源的大语言模型,来帮助对抗机器。
I apologize for the ridiculousness.
我为这荒谬之处道歉。
And like I said, Nathan already knows, I've been a big fan of his for a long time.
就像我说的,内森已经知道,我长期以来一直是他的忠实粉丝。
a big fan of 常用搭配
……的忠实粉丝/非常喜欢……
口语中表达对某人或某事物的强烈喜爱,常用于见面寒暄。
Been a big fan of yours, Sebastian, for a long time.
塞巴斯蒂安,我长期以来一直是你的忠实粉丝。
Been a big fan of yours 常用搭配
一直是你的忠实粉丝
口语中省略主语 I've 的随意说法,用于向对方表达欣赏。
So it's an honor to finally meet you.
所以终于见到你真是荣幸。
it's an honor to finally meet you 句型
终于见到你真是荣幸
it's an honor to finally meet [someone]
正式或半正式场合初次见面时的礼貌表达,表示对见到对方感到荣幸。
Thank you for everything you put out into the world.
感谢你为世界贡献的一切。
put out into the world 常用搭配
向世界发布、贡献(作品等)
用于感谢某人创作或分享的内容,语气真诚、略带正式。
Thank you for the excellent books you're writing.
感谢你正在写的那些优秀的书。
Thank you for teaching us.
感谢你教导我们。
And thank you for talking today. This was fun.
也感谢你今天来谈话。这很有趣。
Thank you for inviting us here and having this human connection, which is extremely valuable human connection.
感谢你邀请我们来到这里,并拥有这种人际联系,这是极其宝贵的人际联系。
Thanks for listening to this conversation with Sebastian Rashka and Nathan Lambert.
感谢收听与塞巴斯蒂安·拉什卡和内森·兰伯特的这次对话。
To support this podcast, please check out our sponsors in the description,
为了支持这个播客,请查看描述中的赞助商,
check out 常用搭配
查看、去看看
口语中常用,表示去看看某物或了解更多信息。
where you can also find links to contact me, ask questions, give feedback, and so on.
在那里你也可以找到联系我、提问、提供反馈等的链接。
And now, let me leave you with some words from Albert Einstein.
现在,让我用阿尔伯特·爱因斯坦的一些话作为结束。
leave you with 常用搭配
用……作为结束/留给你
常用于演讲或节目结尾,表示以某句话或某样东西作为收尾。
It is not that I'm so smart, but I stay with the questions much longer.
并不是我有多聪明,而是我在问题上停留的时间长得多。
Thank you for listening, and hope to see you next time.
感谢收听,希望下次再见。
hope to see you next time 地道口语
希望下次再见
节目或对话结尾常用的告别语,轻松友好。