147. 和蚂蚁灵波沈宇军聊:机器人原生基础模型、大脑和本体的关系、预训练与数据scale up、老师汤晓鸥
沈宇军从视觉生成研究讲到蚂蚁灵波的具身原生路线:为什么机器人大脑需要从物理世界需求出发,重做视觉预训练、实时架构与跨本体数据管线。对谈也给出当前能力的边界——位置随机性已有进步,但机器人的 GPT-1 时刻仍取决于可规模化、可被模型使用的数据。
沈宇军从视觉生成研究讲到蚂蚁灵波的具身原生路线:为什么机器人大脑需要从物理世界需求出发,重做视觉预训练、实时架构与跨本体数据管线。对谈也给出当前能力的边界——位置随机性已有进步,但机器人的 GPT-1 时刻仍取决于可规模化、可被模型使用的数据。
Tommy Wood 从突触连接与修剪出发,解释新奇挑战、clutch 状态、练习节奏和注意力如何共同扩展学习能力。对谈进一步区分有氧、间歇和抗阻训练对脑结构的作用,并讨论营养缺口、痴呆风险与脑震荡恢复的证据边界。
2026 年 4 月,Ai2 后训练负责人 Nathan Lambert 用 6-7 晚跑遍北京、杭州的中国 AI 实验室——阿里巴巴、月之暗面、智谱、美团、清华、小米、蚂蚁百灵、魔搭。这期硅谷 101 的专访是他回美后最完整的一次现场复盘:中国研究员为什么这么年轻、"AGI 展示厅"是什么、开源模型领导权已在 2025 年夏天易手、华为芯片推理可用训练不行、Claude 在中国仍是开发者最爱、以及一个被严重低估的中国短板——数据产业。
Mark Chen 在 Latent Space "cooking series" 一边煮韩式豆腐汤一边讲清楚 OpenAI 的研究路线: scaling laws 在 10 个数量级上没有破,pre-training is dead 是循环出现的伪叙事,evals 行业整体在"裸奔", reasoning(o1)能上线靠的是 Jakub 和 Ilya 的 conviction 而不是显然的算法红利, 3 年路线图的终点是 end-to-end research——模型不仅做实现,还能有 taste。
Gray Swan 两位 CMU 教授创始人讲他们如何把 AI 安全做成一个独立行业。 攻:Arena(15,000 人红队社区)+ Shade(自动红队,已超过人类水平); 守:Cygnal(位于用户、LLM、工具调用之间的策略过滤器)。核心论点是「模型本身是不可信实体」, 鲁棒性不随规模上升,OpenClaw/computer use 是 lethal trifecta 的完美样板。
Ronak Malde 从 Windsurf 一路被 DeepMind 收购,然后掏出收购款回来做 Trajectory.ai—— 一个押注 continual learning 是下一代 AI 产品形态的平台。访谈里他拆了 SDPO 这个自蒸馏 RL 算法、Continuous LoRA 的并发训练栈、以及和 Harvey + Nemotron 3 训出来的法律模型。
一位 MLOps community Amsterdam 的现场分享。讲者把 "context engineering" 拆成一个 很哲学的命题:每周一个新模型,你能控制的只有那个 context window —— 怎么往里塞东西, 比模型本身更重要。从 25% 的经验法则、context 的三分法(deterministic / probabilistic / human),到用大脑结构作类比、把 Karpathy 提议的 "markdown 当 memory" 拓展成带衰减 和重要度评分的 wiki —— 最后以一个 5 分钟硬 timer 的对比 demo 收尾。
田渊栋离开 Meta FAIR 半年后,与 Richard Socher、熊蔡明、Tim Rocktäschel 等 7 位顶级研究员一起创立 Recursive Superintelligence(RSI),6.5 亿美元融资、46.5 亿美元估值。 这一期他第一次系统讲述了 RSI 的技术路线:用 AI 优化 AI(左脚踩右脚),把 auto research 作为商业化第一步,以及他认为现在 AI 能力大约在"满分 10 分的 0.5 分"。 对话还覆盖了 NeoLab 全景、coding 之后的下一波、大厂蒸馏员工的"吸星大法"、以及"水越来越少,鱼必须变成四维生物"的就业大趋势。
两位经济学家——DeepMind 的 Alex Imas 和 Epoch 的 Phil Trammell——讨论 AGI 时代的劳动份额、相关性部门、混乱中段 与再分配难题。他们提出一个反直觉的核心命题:AI 越聪明,它在 GDP 里的相对份额可能反而越小,因为机器经济会内部 闭环;劳动份额能否维持,取决于人类对"人在里头"的内在偏好与变种增长之间的赛跑。
Ethan He 在 xAI 从 0 infra / 0 data / 0 model 起步,3 个月内带几个工程师交付了 Grok Imagine 0.9(首个大规模部署的 audio-video joint generation 模型)。这期访谈把 视频模型的全链路拆开讲: synthetic text-video pair、VAE tokenizer、temporal vs frame-by-frame 压缩、step distillation、长视频的 context 管理。他给 world model 下了一个 三支柱定义 (real time + interactive + long horizon),并提出一个"big claim" —— 现在视觉模型的进步主要来自语言模型,而不是视频模型本身——这也是他刚离开 xAI 想回到 语言模型 / context-aware agent 方向的原因。
张小珺·语言即世界 EP.140,对姚顺宇的 4 小时访谈节选。姚顺宇博士毕业于斯坦福 理论高能物理,2024 年半道出家加入 Anthropic 参与 Claude 3.7、4.5 的强化学习训练; 2025 年 10 月跳槽到 Google DeepMind 做 Gemini 的 ML coding / long horizon。 这期把两家 lab 的打法、coding bet 的内部信号、AI safety 的"幼稚"自我说服、 以及"个人英雄主义时代已经过去了"等小疯言论摊开讲清楚。
Eric Jang 在 sabbatical 期间用大约 10K 美元 + Claude Code 重新实现了 AlphaGo, 在和 Dwarkesh 的对谈里把 MCTS 拆到底——为什么它不是 credit assignment、 为什么它比当下 LLM RL 优雅得多、以及 10 层网络居然能把一个看似 intractable 的搜索问题塞进一次 forward pass。后半段是用 Opus 4.6/4.7 做自动化研究的体感: 超参搜索很强,但还不会"换条路想想看"。
张小珺商业访谈录 #139 期:和俄亥俄州立大学教授、NeoCognition 创始人苏煜做的一次 Agent 技术综述。 从 Logical Agent (1960s-90s) → Neural Agent → Semantic Parsing → Language Agent 的演进史出发, 讨论了 OpenClaw Moment 与 ChatGPT Moment 的相似性、universal digital agent 的目标、 中美科技辐射的不同 pattern,以及 2026 年 Agent 的瓶颈和大厂们的赌注。
Karpathy 在一场炉边对谈里,从"作为程序员从未如此落后"讲起:December 是 agentic 编码工作流真正开始 work 的拐点。他串起 Software 3.0(编程变成 prompting)、可验证性如何造就"锯齿状"智能、vibe coding 与 agentic engineering 的分野,以及人类仍独一无二负责的"理解"。
小米大模型负责人罗福莉的 3.5 小时深度对谈:从春节凌晨 2 点 OpenClaw 觉醒,到 MiMo V2 系列 (Pro / Omni / TTS) 的"悄无声息伏击",再到 Agent 时代后训练算力 1:1、组织扁平化、AGI 两年内可期。 当下范式已从 Chat 切到 Agent —— 1T 基座 + 后训练敏捷性是新的入场券。
Latent Space 与 Moonlake 两位负责人 Chris Manning 与 Fan-yun Sun 的对谈。Moonlake 押的是另一条 world model 路线: 不是更大的视频生成器, 而是 symbolic 推理 + 神经渲染。 Chris 给出唯一的硬定义 — "you only actually have a world model if you can predict, given some action is taken, what is going to change" — 然后顺势公开和 Yann LeCun 撕: "Yann has never appreciated the power of language." Sun 反驳"反 bitter lesson" 的标签, 真正的问题是"what is the right abstraction level today"。Moonlake 内部其实是 两个模型: 推理模型管 causality / persistency, 而 Rie 这个 diffusion model 负责 photorealism — 他们已经把它当作 DLSS 的下一代来卖, "skins for worlds"。
《硅谷101》机器人特辑的开源篇——把机器人 VLA 开源生态拆成四股力量:学院派(OpenVLA、Octo)、巨头生态派(NVIDIA GR00T、Google Gemini Robotics)、创业 + 中国力量(小米、蚂蚁、自变量、清华、OpenMind 等),以及独立成派的 Physical Intelligence(π₀)。逐一剖析这些模型的技术路线、开源动机和阵营关系,最后回到那个核心问题:估值 56 亿美元的 PI 为什么要把核心模型免费给世界?以及一个 Stanford 教授父亲对这场开源运动最朴素的回答。
Mistral 同一周内同时发了 Voxtral TTS、Forge 平台、Leanstral 形式化模型和新的 Mistral Small——这期 Latent Space 让 Pavan Kumar Reddy 和 Guillaume Lample 一次性把这些 发布背后的工程选择讲清楚。Voxtral TTS 是 3B 模型 + 自研 12.5 Hz 神经音频 codec + auto-regressive flow matching head——为了实时流式而不是 SOTA quality 选 AR 路线。 Forge 是把 Mistral 科学团队用了 2 年的 infra 直接给客户:fine-tune 后能"10x cheaper", 并在某些客户项目把一种语言从 0~1% 训到 50% 的 mix。Leanstral 看似是数学家工具, 实际是赌 long-horizon reasoning 的 transfer——Lean 的编译器是天然不可 reward-hack 的判官。最后透露下一代 RL infra 是为"6 hours to get a reward"的 trajectory 设计的。
Alex Lupsasca 是 2024 年 New Horizons in Fundamental Physics Breakthrough Prize(被称为 "Oscar for physics")的得主, 一位黑洞理论物理学家。他追踪 LLM 在科学前沿的能力已经 一年半。GPT-5 发布时 Twitter 反响 "lukewarm" — 但在他的领域, 模型在 30 分钟内复现 了他自己花了很长时间才做出来的好论文。Mark Chen 教了他一个 "priming" 技巧 (先解一道 textbook warmup), GPT-5 就能解决一篇 training-cutoff 之后才发布的论文。 之后, 他和 PhD 导师 Strominger 把一个 32 项之和、卡了一年的 single-minus gluon tree amplitude 问题给了 ChatGPT — 模型在 Strominger 的飞机降落之前就解决了, 还用 作者们都不知道的技巧给出了证明。第二个实验把题目换成 graviton, 模型在一天内吐出 110 页全新的量子引力, 团队用三周验证。这就是 "vibe physics"。
谢赛宁的第一次播客访谈,7 小时马拉松对谈, 覆盖 SJTU ACM Class、UCSD 屠卓文、FAIR 何恺明、NYU 李飞飞, 到 2024 年和 Yann LeCun 共同创立 AMI Labs 的全过程。 贯穿主线:representation learning 是 12 年都没解决的核心问题, LLM 是 virtual intelligence,world model 才是真问题。
Lex Fridman 请来两位"一线做过模型、也写过书"的研究者做 2026 年初的 AI state-of-the-art 盘点: Sebastian Raschka 从 GPT-2 一路手撕到 Qwen3 / Gemma 3, 最擅长从架构里读故事;Nathan Lambert 是 AI2 研究员、RLHF 书作者、atom 项目 发起人,frontier 与 open-source 两边都站过。两人聊了 DeepSeek 时刻、Opus 4.5 神话、RLVR 的"假 aha"、scaling 的三个轴、AI 2027 的时间线推后、Anthropic $1.5B 和解、CUDA 的真护城河、atom project,一直到 100 年后世界的样子。
田渊栋在 Meta 10 月 600 人 AI 裁员风波中被裁,被裁前已有 offer。在这场访谈里,他坦诚谈了对 LLM 路线、Scaling Law、RL 与 SFT、AI 人才市场的判断,回顾了在 FAIR 十年最大的收获——research taste,并谈到他想成为"超级研究员"的下一步设想。
当世最强的数学家谈他亲手用过 AI 之后的判断:想法生成的成本几乎归零, 但瓶颈搬到了 verification 这一侧;50 道 Erdős 问题被 AI 攻克之后随即陷入瓶颈; AI 像在黑暗中乱跳的机器人,能跨过低墙、爬不上悬崖;breadth × depth 才是 数学的下一步——但要先重设整个学术工作流。
Andrej Karpathy 在 Dwarkesh Podcast 的长访谈。他给出一份冷静的"祛魅":这是 智能体的十年而非元年;我们造的不是动物而是"幽灵"——通过模仿互联网而来的数字 实体。他剖析了 RL 的根本缺陷("用吸管吸取监督信号")、模型坍缩、自动驾驶式的 "九分进军",以及他为何离开前沿实验室转去做教育。
硅谷101 陈茜 复盘 GPT-5 发布会的「就这?!」翻车现场——SWE-bench 图表错乱、 伯努利效应解错、4o 用户集体抗议——并对谈多位 AI 技术专家解释 GPT-5 的真相: 它不是端到端超级大模型,而是路由器拼接方案,业内戏称「GPT-4.99」。Scaling Law 确实碰壁,下一步突破要靠 RL + universal verifier、多模态/世界模型、或者跳出 Transformer 的 JEPA 等新架构。
Karpathy 在 YC AI Startup School 的演讲:软件 70 年没怎么变,却在最近几年被快速改写了两次—— 从 Software 1.0(代码)到 2.0(神经网络权重)再到 3.0(英文 prompt)。他用一连串类比拆解 LLM: 它像电网、像晶圆厂,但最像停留在 1960 年代的操作系统;它是有一身认知缺陷的"人类幽灵"。 最后落到怎么和它一起工作——部分自治应用、自治滑块、生成-验证循环,以及为 Agent 重写基础设施。
硅谷101 走进硅谷一家量子计算实验室,亲手帮忙装机。围绕谷歌 Willow 的"低于阈值"突破、五家巨头(IBM/Google/Amazon/Nvidia/Microsoft)截然不同的技术路线、加密货币和 AI 受到的冲击,以及最现实的金融与材料应用,给出一个不靠口号、靠路线图的判断:**业内大多数人会认同的实用化时间点是 20 年左右**。
Lex Fridman 与 Anthropic 的三位核心人物同场对谈: CEO Dario Amodei 讲 scaling 假设、RSP 的 if-then 结构、"race to the top" 战略与 Machines of Loving Grace; character lead Amanda Askell 讲 Claude 的性格工程、sycophancy 与最优失败率; interpretability 共同创始人 Chris Olah 讲 features、circuits、superposition 和那个著名的 deception feature。三人从战略、产品、研究三个层面拼出 Anthropic 对 "AI inside" 的完整 stereo view。
Emotion scientist Dr. Dacher Keltner explains the neuroscience of awe — how shifting perception from "small to vast" triggers measurable health benefits including reduced inflammation, less pain, and better brain health years later. The conversation spans the expanded taxonomy of 20 facial expressions, the role of teasing and embarrassment in group bonding, the loneliness epidemic, music as the fastest path to collective consciousness, and practical awe design for cities and daily life.
Dr. Marc Breedlove explains how prenatal testosterone shapes sexual orientation, the fraternal birth order effect, gay rams, brain sexual dimorphisms, and why sexual orientation is not a choice — through the lens of decades of biological research.
Dr. Rhonda Patrick joins Andrew Huberman for a comprehensive deep-dive into evidence-based health protocols covering exercise (VILPA, strength training), gut-inflammation cascades, omega-3 as her top supplement, cortisol management, intermittent fasting mechanics, creatine for brain health, and her full supplement stack with dosages and rationale.
Dr. Richie Davidson, a pioneer in meditation neuroscience, explains how just 5 minutes of daily meditation for 30 days produces measurable reductions in depression, anxiety, and inflammation. He introduces the concept of "the lactate of the mind" — the initial anxiety of meditation as an adaptation signal — and presents his four-pillar framework for human flourishing: awareness, connection, insight, and purpose.
Neil Turok 在这期 Theories of Everything 长访谈里提出:量子引力或许根本不需要弦、额外维度或多元宇宙。 他和博士生 Sam Bateman 把 1977 年 Stelle 的二次引力理论,配合一个对玻恩规则的微小推广(让概率定义在 Klein 空间而非希尔伯特空间里), 绕过了高阶导数理论的两个 175 年老咒语:奥斯特罗格拉茨基不稳定性和负范数"幽灵态"。
Professor Renato Renner of ETH Zurich explains his no-go theorem showing quantum theory contradicts itself when applied to observers who are themselves quantum systems. The conversation covers three incompatible assumptions (universality, consistency, single outcomes), a gravity-based escape route involving reference frame loops, the black hole information paradox, and why choosing an interpretation of quantum mechanics is ultimately an emotional decision.
Harvey Friedman — the youngest professor in recorded history (a Stanford appointment at 18) and the author of the last paper Kurt Gödel sponsored for PNAS — gives his first podcast. The conversation moves from Gödel's two often-conflated incompleteness theorems to Friedman's 60-year program to push incompleteness out of set-theoretic exile and into the kind of finite, combinatorial mathematics that working mathematicians cannot dismiss. Along the way: TREE(3), the divine consistency proof, embedded maximality, and a quiet meditation on AI as a form of immortality.
Michael Nielsen and Dwarkesh Patel explore how scientific progress actually happens — from the messy reality of falsification to hostile verification loops that mislead for decades. They discuss why the tech tree is far vaster than we realize, why alien civilizations would develop radically different technologies, and what this means for AI-accelerated science.
Dr. Jenny Wagner explains why gravitational lensing data only constrains local properties of mass distributions, making every grand dark matter map a model-driven extrapolation. She argues the inverse problem approach — reasoning from data to necessary models — could reshape cosmology and the scientific method itself.
Professor JB Manchak proves that no amount of empirical data — even from every point in the universe — can determine its global structure. He introduces Heraclitus spacetimes (maximally asymmetric universes where local structure determines global structure), and draws surprising parallels between cosmic underdetermination and Zen Buddhist non-self.
历史学家 Lars Brownworth 用 1.5 小时把 793–1066 年的 Viking 时代拆给你看: 从 Lindisfarne 那场让欧洲修士相信"世界末日来了"的袭击,到长船 70–120 mi/day 的恐怖速度, 再到 Ragnar、Valhalla、Berserkers、Leif Erikson 的 Vinland、东到 Constantinople 的 Varangian Guard。最后落到一个反差:这个 destroyer 民族只用了三代人就变成 builder, 造出了 Normandy、Kievan Rus 与现代欧洲的雏形。