<?xml version="1.0" encoding="UTF-8"?>
<feed xmlns="http://www.w3.org/2005/Atom">
  <title>Mochiao Chen</title>
  <subtitle>金融、计算与写作。Mochiao Chen 的个人主页、博客与独立工具集。</subtitle>
  <link href="https://mochiaochen.github.io/feed.xml" rel="self" type="application/atom+xml"/>
  <link href="https://mochiaochen.github.io/" rel="alternate" type="text/html"/>
  <id>https://mochiaochen.github.io/</id>
  <updated>2026-10-06T00:53:08+08:00</updated>
  <author><name>Mochiao Chen</name></author>
  
  <entry>
    <title>杰文斯悖论</title>
    <link href="https://mochiaochen.github.io/writing/2026/10/jevons-paradox/" rel="alternate" type="text/html"/>
    <published>2026-10-06T00:30:00+08:00</published>
    <updated>2026-10-06T00:30:00+08:00</updated>
    <id>https://mochiaochen.github.io/writing/2026/10/jevons-paradox</id>
    <content type="html" xml:base="https://mochiaochen.github.io/writing/2026/10/jevons-paradox/">&lt;h2 id=&quot;一问题的提出&quot;&gt;一、问题的提出&lt;/h2&gt;

&lt;p&gt;杰文斯悖论（Jevons Paradox）指的是这样一个现象：&lt;strong&gt;某种资源的使用效率提高之后，该资源的总消费量反而上升&lt;/strong&gt;。&lt;/p&gt;

&lt;p&gt;这个命题出自英国经济学家 William Stanley Jevons 于 1865 年出版的《The Coal Question》。当时英国正在担忧本国煤炭储量的耗竭会终结其工业优势，主流观点认为可以通过提高蒸汽机的燃料效率来延长煤炭的可用年限。Jevons 认为这个推论完全颠倒了因果。他观察到的历史事实是：从 Newcomen 引擎到 Watt 的分离式冷凝器，蒸汽机的单位功率煤耗显著下降，而英国的煤炭消费量在同一时期不但没有下降，反而以更快的速度增长。&lt;/p&gt;

&lt;p&gt;他的解释是：效率提升降低了「使用蒸汽动力」这项服务的实际成本，从而使蒸汽机在此前不经济的场合（深层矿井排水、铁路、远洋航运、以及各类原本用水力或畜力的工业环节）变得有利可图。技术的适用边界扩张了，新的用途被创造出来，需求的扩张幅度超过了单位需求上节省的燃料。&lt;/p&gt;

&lt;p&gt;需要强调的是，Jevons 的论证对象是&lt;strong&gt;总量&lt;/strong&gt;，不是单位效率。单位煤耗确实下降了，这一点从未有争议。争议在于总量的方向。&lt;/p&gt;

&lt;h2 id=&quot;二形式化表述&quot;&gt;二、形式化表述&lt;/h2&gt;

&lt;p&gt;现代文献用「回弹效应」（rebound effect）这个框架来处理这个问题，杰文斯悖论是其中的一个极端情形。&lt;/p&gt;

&lt;p&gt;设能源服务的产出为 $S$（例如「有效照明流明小时」「吨公里运输量」「室内维持在 20 ℃ 的小时数」），能源投入为 $E$，能源效率定义为&lt;/p&gt;

&lt;p&gt;[\varepsilon = \frac{S}{E}]&lt;/p&gt;

&lt;p&gt;于是 $E = S/\varepsilon$。取对数并对 $\ln \varepsilon$ 求导：&lt;/p&gt;

&lt;p&gt;[\frac{d\ln E}{d\ln \varepsilon} = \frac{d\ln S}{d\ln \varepsilon} - 1]&lt;/p&gt;

&lt;p&gt;工程学式的直觉隐含假设了 $d\ln S / d\ln\varepsilon = 0$，即服务需求量固定不变，因此效率提高 1% 就节能 1%。这个假设正是问题所在。&lt;/p&gt;

&lt;p&gt;能源服务的有效价格是 $p_S = p_E / \varepsilon$。在能源价格 $p_E$ 外生给定时，效率提高等价于服务价格等比例下降。记能源服务需求的自价格弹性为&lt;/p&gt;

&lt;p&gt;[\eta = \frac{d\ln S}{d\ln p_S} &amp;lt; 0]&lt;/p&gt;

&lt;p&gt;代入可得&lt;/p&gt;

&lt;p&gt;[\frac{d\ln E}{d\ln \varepsilon} = -\eta - 1]&lt;/p&gt;

&lt;p&gt;定义回弹率 $R$ 为「被需求扩张吃掉的那部分预期节能」占预期节能的比例，则&lt;/p&gt;

&lt;table&gt;
  &lt;tbody&gt;
    &lt;tr&gt;
      &lt;td&gt;[R = 1 + \frac{d\ln E}{d\ln\varepsilon} = -\eta =&lt;/td&gt;
      &lt;td&gt;\eta&lt;/td&gt;
      &lt;td&gt;]&lt;/td&gt;
    &lt;/tr&gt;
  &lt;/tbody&gt;
&lt;/table&gt;

&lt;p&gt;这个简洁的结果给出了完整的分类：&lt;/p&gt;

&lt;table&gt;
  &lt;thead&gt;
    &lt;tr&gt;
      &lt;th&gt;$\eta$ 的取值&lt;/th&gt;
      &lt;th&gt;回弹率 $R$&lt;/th&gt;
      &lt;th&gt;结果&lt;/th&gt;
    &lt;/tr&gt;
  &lt;/thead&gt;
  &lt;tbody&gt;
    &lt;tr&gt;
      &lt;td&gt;$\eta = 0$&lt;/td&gt;
      &lt;td&gt;$0$&lt;/td&gt;
      &lt;td&gt;零回弹，节能全额兑现&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;$-1 &amp;lt; \eta &amp;lt; 0$&lt;/td&gt;
      &lt;td&gt;$0 &amp;lt; R &amp;lt; 1$&lt;/td&gt;
      &lt;td&gt;部分回弹，仍然节能但少于预期&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;$\eta = -1$&lt;/td&gt;
      &lt;td&gt;$R = 1$&lt;/td&gt;
      &lt;td&gt;完全抵消，总量不变&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;$\eta &amp;lt; -1$&lt;/td&gt;
      &lt;td&gt;$R &amp;gt; 1$&lt;/td&gt;
      &lt;td&gt;&lt;strong&gt;反弹超射（backfire），即杰文斯悖论&lt;/strong&gt;&lt;/td&gt;
    &lt;/tr&gt;
  &lt;/tbody&gt;
&lt;/table&gt;

&lt;p&gt;所以在上述能源价格外生、其他成本不变的局部均衡设定下，杰文斯悖论在形式上等价于一个条件：&lt;strong&gt;能源服务的需求价格弹性在绝对值上大于 1&lt;/strong&gt;。&lt;/p&gt;

&lt;p&gt;这里有一个技术性的注意事项。Sorrell &amp;amp; Dimitropoulos（2008, &lt;em&gt;Ecological Economics&lt;/em&gt;）指出，严格来说应当区分「效率弹性」与「价格弹性」。二者不等价，因为效率提升通常伴随着设备资本成本的上升，而且消费者对「效率标签」的行为反应未必与对「电价下降」的反应相同。用价格弹性作为代理变量，倾向于高估回弹。&lt;/p&gt;

&lt;h2 id=&quot;三传导机制的分解&quot;&gt;三、传导机制的分解&lt;/h2&gt;

&lt;p&gt;上面的推导过于压缩，实际的传导路径需要分层理解。文献中的标准分类如下。&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;直接回弹（direct rebound）&lt;/strong&gt;，作用于同一项能源服务内部，可进一步拆为两个新古典效应：&lt;/p&gt;

&lt;ul&gt;
  &lt;li&gt;替代效应：能源服务相对于其他商品变便宜了，消费者用它替代别的东西。空调更省电，于是开得更久、温度设得更低。&lt;/li&gt;
  &lt;li&gt;收入效应：账单下降释放出的实际购买力，其中一部分会重新流向该能源服务本身。&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;间接回弹（indirect rebound）&lt;/strong&gt;，节省下来的支出流向了其他商品和服务，而这些商品的生产与消费本身也内含能源。省下的电费用来买机票，这部分排放需要计入。此外还有内含能源（embodied energy）问题：更高效的设备本身在制造环节可能更耗能。&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;经济体层面的回弹（economy-wide rebound）&lt;/strong&gt;，这是 Jevons 本人真正强调的层次，也是最难量化的层次，包含几条渠道：&lt;/p&gt;

&lt;ol&gt;
  &lt;li&gt;&lt;strong&gt;能源价格的一般均衡反馈&lt;/strong&gt;。效率提升压低了对能源的总需求，在供给曲线不完全弹性时，能源市场出清价格下降，这又反过来刺激所有部门的能源使用。这条渠道在化石燃料这类全球定价的商品上尤其重要，构成了所谓的「绿色悖论」的近亲。&lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;产出效应与要素替代&lt;/strong&gt;。对企业而言，能源效率提升等价于一次成本降低，即生产函数的正向技术冲击。企业会扩大产出规模，同时可能用能源替代资本或劳动。&lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;增长效应&lt;/strong&gt;。这是最长期的渠道。能源效率的持续提升本身就是全要素生产率增长的组成部分，它抬高了资本回报率，加速资本积累，从而扩张整个经济的规模。Saunders（1992, &lt;em&gt;The Energy Journal&lt;/em&gt;）在新古典增长模型中证明：若生产函数为 Cobb-Douglas 且能源供给完全弹性，则能源效率提升&lt;strong&gt;必然&lt;/strong&gt;导致 backfire。在更一般的 CES 设定下，是否 backfire 取决于替代弹性与能源要素份额之间的关系（Saunders, 2008）。&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Khazzoom（1980）从家电的直接回弹入手，Brookes（1990）从宏观增长入手，二者的结论后来被 Saunders 合称为 &lt;strong&gt;Khazzoom-Brookes 假说&lt;/strong&gt;：在经济体层面，能源效率的提升会提高而非降低能源消费总量。&lt;/p&gt;

&lt;h2 id=&quot;四经验证据&quot;&gt;四、经验证据&lt;/h2&gt;

&lt;p&gt;这是整个议题上分歧最大的部分。大致可以说：直接回弹的证据相当扎实且数值温和，经济体层面的回弹证据薄弱且数值分散。&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;直接回弹&lt;/strong&gt;。Greening、Greene &amp;amp; Difiglio（2000, &lt;em&gt;Energy Policy&lt;/em&gt;）与 Sorrell（2007, UKERC 综述）汇总的估计值大体落在：家用取暖与制冷 10%–30%，私人小汽车出行 10%–30%，家用电器 0%–20%。Small &amp;amp; Van Dender（2007, &lt;em&gt;The Energy Journal&lt;/em&gt;）对美国的研究发现，交通领域的回弹率随收入上升而系统性下降，在全样本平均条件下，短期与长期回弹估计分别为 4.5% 和 22.2%。这符合直觉：当一项服务已经接近需求饱和时，再便宜也不会消费更多。发达国家的室内照明和室温就属于这一类，弹性接近零。&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;历史长周期证据&lt;/strong&gt;。照明是最经典的案例。Fouquet &amp;amp; Pearson（2006, 2012）重建了英国 1300 年至 2000 年的照明服务价格与消费序列，发现单位流明成本下降了数千倍，而人均照明消费的增长幅度更大，即在工业革命至 20 世纪中期的大部分时段内，照明的需求价格弹性在绝对值上超过 1，属于典型的杰文斯情形。但他们同时发现，这个弹性在 20 世纪后期急剧衰减至接近零，因为照明需求已经饱和。Tsao 等人（2010, &lt;em&gt;J. Phys. D&lt;/em&gt;）曾据此讨论 LED 普及可能带来的照明需求扩张。能耗的最终方向仍取决于需求饱和程度与实际回弹，不能仅凭单位效率的改善判断。&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;经济体层面&lt;/strong&gt;。Gillingham、Rapson &amp;amp; Wagner（2016, &lt;em&gt;Review of Environmental Economics and Policy&lt;/em&gt;）综述后认为，总回弹的合理上限大约为 60%，而多数研究指向更小的效应，backfire 是可能的但缺乏证据支持，因此不应作为政策制定的基准假设。相反，Brockway 等人（2021, &lt;em&gt;Renewable and Sustainable Energy Reviews&lt;/em&gt;）复核了 33 项研究，认为经济体层面的回弹可能侵蚀超过一半的预期节能，并明确指出现有证据无法排除 backfire，主张 IPCC 等机构的减排路径可能系统性高估了能效改进的贡献。&lt;/p&gt;

&lt;p&gt;这两篇立场对立的综述是理解当前学术分歧的最佳入口。分歧的根源在于：经济体层面的回弹需要一个反事实的「若无效率提升的世界」，而这个反事实只能靠 CGE 模型或增长核算来构造，其结果对生产函数形式、替代弹性、能源供给弹性的设定高度敏感。这使结果的检验与跨研究比较变得困难，需要清楚报告模型假设和敏感性分析。&lt;/p&gt;

&lt;h2 id=&quot;五什么时候更容易出现悖论&quot;&gt;五、什么时候更容易出现悖论&lt;/h2&gt;

&lt;p&gt;综合上述机制，可以列出几个提高 backfire 概率的条件：&lt;/p&gt;

&lt;ul&gt;
  &lt;li&gt;&lt;strong&gt;市场远未饱和&lt;/strong&gt;，存在大量被高成本抑制的潜在用途。这解释了为何 19 世纪的煤炭和当代发展中国家的电力更容易出现悖论，而发达国家的家庭取暖不会。&lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;能源成本在该活动总成本中占比高&lt;/strong&gt;。占比越高，效率提升带来的相对价格冲击越大。铝冶炼、氨合成、数据中心的算力属于这一类；而对于一个白领办公室，电费占比微不足道，回弹自然接近零。&lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;该技术是通用目的技术（general purpose technology）&lt;/strong&gt;。蒸汽机、电力、内燃机、以及当下的计算，其效率提升会向全经济扩散并催生全新的应用形态，产出效应极强。计算领域的 Koomey 定律与算力消费的爆炸式增长是一个正在发生的样本。&lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;能源供给曲线较为平坦&lt;/strong&gt;，价格反馈渠道弱化了抑制作用。&lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;时间尺度长&lt;/strong&gt;。直接回弹在数月内实现，产出与增长效应需要数十年，因此短期实证研究会系统性低估总回弹。&lt;/li&gt;
&lt;/ul&gt;

&lt;h2 id=&quot;六常见的误读&quot;&gt;六、常见的误读&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;误读一：把杰文斯悖论等同于回弹效应。&lt;/strong&gt; 回弹效应几乎必然存在且几乎必然为正，这是标准的价格反应，没有任何悖论意味。杰文斯悖论是 $R &amp;gt; 1$ 的特例。绝大多数实证研究支持的是 $0 &amp;lt; R &amp;lt; 1$。把「回弹存在」表述为「杰文斯悖论成立」是一个常见的量级错误。&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;误读二：以此论证能效政策无效。&lt;/strong&gt; 这里混淆了两个不同的目标函数。即便在 backfire 的情形下，效率提升仍可能增加社会福利：人们获得了更多的照明、更多的出行、更多的工业产出，成本更低。能源消费量上升的原因恰恰是效率提升创造了真实的价值。用能源消费量作为唯一的评价指标，等于把手段当成了目的。&lt;/p&gt;

&lt;p&gt;正确的政策推论是另一个：&lt;strong&gt;能效政策无法单独承担减排任务，因为单位效率的改善本身不能保证总能源消费或总排放下降&lt;/strong&gt;。回弹的全部传导渠道都以「有效价格下降」为枢纽，因此对策也在价格端：碳定价或总量交易可以在效率提升释放需求的同时抬高单位排放的成本，从而把回弹锁在闸门之内。总量控制机制（cap-and-trade）在逻辑上尤其干净，因为它直接对总量设限，使回弹只能表现为配额价格的变化而非排放的增加。能效标准与碳定价是互补而非替代的关系。&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;误读三：把它当作技术乐观主义或增长主义的论据。&lt;/strong&gt; 值得注意的是，这个概念在当代被两个对立阵营同时引用。反对气候管制的一方用它论证「技术进步会自然解决问题，管制多余」；degrowth 阵营则用它论证「技术进步永远无法实现增长与排放的脱钩，必须直接限制经济规模」。两种用法都超出了这个命题本身的实证支撑范围。杰文斯悖论只是一个关于价格弹性与一般均衡反馈的观察，它不自动蕴含任何一方的规范结论。&lt;/p&gt;

&lt;h2 id=&quot;七一个更抽象的概括&quot;&gt;七、一个更抽象的概括&lt;/h2&gt;

&lt;p&gt;如果要把这个悖论的内核抽离出来，它其实是&lt;strong&gt;对「其他条件不变」这一假设的一次反例演示&lt;/strong&gt;。工程学的节能测算天然是局部均衡的：固定使用模式，改进设备，计算差额。经济学告诉你，使用模式本身是价格的函数，而价格正是你所改变的那个东西。一个改变了约束条件的干预，无法用「约束条件不变」的框架来评估其后果。&lt;/p&gt;

&lt;p&gt;这个结构在能源之外反复出现。拓宽道路以缓解拥堵，结果诱发了新的出行需求（Duranton &amp;amp; Turner, 2011, &lt;em&gt;AER&lt;/em&gt; 的「道路交通基本定律」，其估计的诱发需求弹性接近 1）；提高存储容量以整理文件，结果堆积了更多文件；抗生素的高效使得处方门槛降低，从而加速了耐药性的演化。这些都是同一个逻辑：&lt;strong&gt;效率是一种价格削减，而价格削减会被需求吸收&lt;/strong&gt;。&lt;/p&gt;

&lt;p&gt;需要吸收多少，取决于弹性。而弹性取决于需求是否已经饱和。&lt;/p&gt;

&lt;hr /&gt;

&lt;h2 id=&quot;参考资料&quot;&gt;参考资料&lt;/h2&gt;

&lt;ul&gt;
  &lt;li&gt;Small &amp;amp; Van Dender (2007), &lt;a href=&quot;https://its.uci.edu/research_products/published-journal-article-fuel-efficiency-and-motor-vehicle-travel-the-declining-rebound-effect/&quot;&gt;Fuel Efficiency and Motor Vehicle Travel: The Declining Rebound Effect&lt;/a&gt;.&lt;/li&gt;
  &lt;li&gt;Gillingham, Rapson &amp;amp; Wagner (2016), &lt;a href=&quot;https://gwagner.com/rebound&quot;&gt;The Rebound Effect and Energy Efficiency Policy&lt;/a&gt;.&lt;/li&gt;
  &lt;li&gt;Brockway et al. (2021), &lt;a href=&quot;https://eprints.whiterose.ac.uk/id/eprint/171952/&quot;&gt;Energy efficiency and economy-wide rebound effects&lt;/a&gt;.&lt;/li&gt;
&lt;/ul&gt;
</content>
  </entry>
  
  <entry>
    <title>免于必然之后</title>
    <link href="https://mochiaochen.github.io/writing/2026/10/after-necessity/" rel="alternate" type="text/html"/>
    <published>2026-10-06T00:30:00+08:00</published>
    <updated>2026-10-06T00:30:00+08:00</updated>
    <id>https://mochiaochen.github.io/writing/2026/10/after-necessity</id>
    <content type="html" xml:base="https://mochiaochen.github.io/writing/2026/10/after-necessity/">&lt;p&gt;&lt;em&gt;AGI 之后的人类生活：一份思想史的清点、批判与推演&lt;/em&gt;&lt;/p&gt;

&lt;hr /&gt;

&lt;h2 id=&quot;导言问题的形状&quot;&gt;导言：问题的形状&lt;/h2&gt;

&lt;p&gt;「AGI 之后人类过什么样的生活」是一个被提问频率远高于被认真处理频率的问题。它在公共讨论中通常以两种失效的形式出现：一种是技术乐观主义的许诺，说人类终于可以去追求艺术、哲学与爱；另一种是反乌托邦的警告，说人类将沦为多余的、被供养的、无所事事的存在。这两种回答共享同一个缺陷，就是它们都把「人不必劳动」当作一个自明的状态，而没有追问这个状态内部的结构。&lt;/p&gt;

&lt;p&gt;要把这个问题问对，第一步是把它拆开。它至少包含三个层次不同、难度递增、解决路径完全不重合的子问题。&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;第一层是分配问题。&lt;/strong&gt; 当劳动不再是价值创造的主要投入，产出归谁所有？谁有权决定 AGI 的算力被用来做什么？这是政治经济学问题。它有明确的、可操作的解决方案模板：累进资本税、公共所有权、主权财富基金、全民分红、劳动份额补贴。这一层的困难在于政治可行性与权力配置，而不在于知识匮乏。人类已经知道该怎么做，问题是有没有人有能力让它发生。&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;第二层是目的问题。&lt;/strong&gt; 当人不必工作，人的时间如何被组织？身份从哪里来？共同体如何维系？这是社会学与心理学问题。它有部分的经验证据（失业研究、退休研究、有闲阶级史、基本收入实验），但证据的外推能力有限，因为历史上从未有过全社会同时进入这一状态的先例。&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;第三层是意义问题。&lt;/strong&gt; 当机器在几乎所有维度上做得比人好、比人快、比人可靠，人的行为还有没有客观价值？如果我写的一首诗、做的一次证明、给出的一个诊断，都有一个更好的机器版本随时可得，我做这件事的理由还剩下什么？这是形而上学与伦理学问题。它没有工程解，也没有政策解。它只可能有一个概念上的重新安排，或者一个失败。&lt;/p&gt;

&lt;p&gt;这三层之间的关系不是并列的。&lt;strong&gt;分配问题在时间上优先&lt;/strong&gt;，因为它决定了后两个问题会以什么样的社会形态被提出；&lt;strong&gt;意义问题在难度上优先&lt;/strong&gt;，因为它是唯一一个无法通过任何再分配机制解决的问题。&lt;u&gt;一个分配得极其公平的世界，仍然可能是一个人人都觉得自己多余的世界。&lt;/u&gt;&lt;/p&gt;

&lt;p&gt;思想史上的哲人几乎从未讨论过 AGI。这是显然的，也是必须承认的。但他们高强度地、反复地讨论过后两个问题，原因很简单：&lt;strong&gt;「一部分人不必劳动」在历史上并非假想，而是长期存在的实存制度。&lt;/strong&gt; 古典世界的公民依托奴隶制获得闲暇，近代欧洲的食利者与绅士阶层依托地租与资本获得闲暇，当代的退休人群与继承财富者依托社会保险与积累获得闲暇。这些人群的生活状态、心理结果与文化产出，构成了人类关于「免于必然之后会发生什么」的唯一自然实验。&lt;/p&gt;

&lt;p&gt;AGI 相对于这些先例，只增加了两个真正新的变量。第一，它把这种状态从少数人扩展到全体，从而消除了「有闲」原本包含的相对性与地位含义。第二，也是更根本的，它第一次把「人在能力上被全面超越」加了进来。历史上的有闲阶级不必劳动，但他们仍然是自己所在领域中最好的那批人；他们的诗、他们的科学、他们的政治判断，是当时世界上能得到的最好的东西。AGI 之后的人不是这样。&lt;/p&gt;

&lt;p&gt;这篇文章的目的不是预测。思想史无法预测，任何声称能从柏拉图推出 2050 年社会结构的论证都是骗人的。它能做的是提供概念工具：把一个混乱的、被情绪主导的问题，分解成若干个有明确论证结构、有反驳、有经验检验方向的子命题。以下的清点大致按时间顺序，但每一部分都会指出该论证在今天的适用性与失效边界。&lt;/p&gt;

&lt;hr /&gt;

&lt;h2 id=&quot;第一部-前史人类曾经如何想象免于劳动&quot;&gt;第一部 前史：人类曾经如何想象免于劳动&lt;/h2&gt;

&lt;h3 id=&quot;一亚里士多德的自动梭子&quot;&gt;一、亚里士多德的自动梭子&lt;/h3&gt;

&lt;p&gt;关于自动化的最早文本，也是最惊人的一段，出现在《政治学》第一卷论奴隶制的语境里（1253b33–1254a1）。亚里士多德在那里设想了一种反事实：如果每件工具都能听从命令、甚至预先感知而自行完成自己的工作，就像传说中代达罗斯的雕像会自己走动、赫菲斯托斯的三足鼎会自己驶向众神的集会那样；如果梭子能自己织布，拨子能自己弹琴，那么工匠就不再需要助手，主人也不再需要奴隶。&lt;/p&gt;

&lt;p&gt;必须注意亚里士多德提出这个设想的目的。他提出它是为了否定它。这个反事实在他看来显然不成立，因此奴隶制作为家政的必要组成部分被保住了。但论证的结构留了下来，而且这个结构比它所服务的结论重要得多：&lt;mark&gt;&lt;strong&gt;自动机械与奴役在功能上是可替代的。&lt;/strong&gt;&lt;/mark&gt; 一个社会要让一部分人获得闲暇，要么让另一部分人替他们劳动，要么让机器替他们劳动。这两条路在逻辑上是同一条路的两个分支。&lt;/p&gt;

&lt;p&gt;这一洞察在思想史上被反复捡起。马克思在《资本论》第一卷论机器与大工业的章节中专门引用了亚里士多德的这段话，用来嘲讽资产阶级经济学家对机器解放作用的乐观。Norbert Wiener 在 1950 年的《人有人的用处》（&lt;em&gt;The Human Use of Human Beings&lt;/em&gt;）中再次引用，并给出了一个远为尖锐的推论：自动化机器是奴隶劳动的经济等价物，因此任何接受与机器劳动竞争的人类劳动，都将被压低到奴隶劳动的经济条件之下。Wiener 的意思并非工人会变成奴隶，他说的是，在纯市场逻辑下，人类劳动的价格将被机器的边际成本定义，而这个成本正在趋近于零。这个论证在 1950 年是推测，在今天已经可以被部分检验。&lt;/p&gt;

&lt;h3 id=&quot;二σχολή闲暇是原初劳动是缺失&quot;&gt;二、σχολή：闲暇是原初，劳动是缺失&lt;/h3&gt;

&lt;p&gt;比自动梭子更深的洞察藏在希腊语的词源里。「闲暇」是 σχολή（scholē），这个词后来经拉丁语 schola 演变为英语的 school，也就是说，学校在词源上就是「闲暇的场所」。而「工作、忙碌」是 ἀσχολία（ascholia），字面意思是「无闲暇」，前缀 ἀ- 是否定。&lt;/p&gt;

&lt;p&gt;这个构词关系本身就是一种哲学主张，而且是与现代人的直觉完全相反的主张。在希腊人的语言结构里，闲暇是基本状态，劳动是这个基本状态的缺失或中断。在现代人的语言结构里，情况正好颠倒：工作是基本状态，闲暇是工作之间的间隙，是 leisure time、是 off-hours、是「下班之后」。我们用工作来定义休息，希腊人用闲暇来定义工作。&lt;/p&gt;

&lt;p&gt;Josef Pieper 在《闲暇：文化的基础》（&lt;em&gt;Muße und Kult&lt;/em&gt;, 1948）中把这个颠倒当作现代性诊断的核心：一个把自身理解为「全面工作的世界」（Welt der totalen Arbeit）的文明，已经在语言层面丧失了理解闲暇的能力，因而它的闲暇必然只能被理解为「为了更好地工作而进行的恢复」，也就是被工作彻底殖民的闲暇。这个诊断对 AGI 讨论极为关键，因为它意味着：&lt;strong&gt;当代人在设想「不必工作的生活」时，使用的是一套已经预设了工作中心地位的概念工具，这套工具本身可能不足以设想它试图设想的对象。&lt;/strong&gt;&lt;/p&gt;

&lt;h3 id=&quot;三沉思生活与政治生活一个未解决的张力&quot;&gt;三、沉思生活与政治生活：一个未解决的张力&lt;/h3&gt;

&lt;p&gt;亚里士多德对「闲暇里该做什么」给出了两个不完全兼容的答案，这个张力一直延续到今天。&lt;/p&gt;

&lt;p&gt;在《政治学》第七、八卷，他的答案偏向教育与公民生活。他在那里提出了一个至今没有被超越的论点：&lt;strong&gt;闲暇有它自己的德性要求，闲暇比劳动更难。&lt;/strong&gt; 他以斯巴达为例（1334a），指出斯巴达式的城邦在战争与必需的压力下表现优异，却在和平与闲暇到来时迅速腐败瓦解，因为他们的整套教育只训练了应对困难的能力，从未训练过应对无困难状态的能力。斯巴达人「只要在打仗就能保全自己，一旦获得了帝国就毁灭了自己，因为他们不知道如何使用闲暇」。第八卷论教育（1337b–1338a）随即论证，音乐之所以必须被纳入自由人的教育，恰恰因为它没有实用功能：它训练的是在闲暇中做值得做的事的能力。&lt;/p&gt;

&lt;p&gt;在《尼各马可伦理学》第十卷，答案则偏向 θεωρία（沉思）。亚里士多德在那里论证最高的幸福是沉思活动，理由是它最自足（沉思者比任何人都更少依赖他人）、最持久、最不为别的目的服务、最接近神性的活动方式。这是一个把好生活从公共领域抽回到个体心智的方案。&lt;/p&gt;

&lt;p&gt;这个张力在 AGI 语境下变得极其尖锐，原因是：&lt;strong&gt;沉思路径直接暴露在 AGI 的能力威胁之下，公民行动路径则相对免疫。&lt;/strong&gt; 如果最高的生活是理解世界，而机器理解得比我好，那么这条路被堵死了；但如果最高的生活是与他人共同决定我们如何生活，那么这条路在原则上不受威胁，因为「我们如何生活」这个问题的主语无法被外包。Arendt 在 1958 年选择了第二条路，理由正在于此，我们后面会详细讨论。&lt;/p&gt;

&lt;h3 id=&quot;四安乐乡民间想象中的丰饶&quot;&gt;四、安乐乡：民间想象中的丰饶&lt;/h3&gt;

&lt;p&gt;哲学传统之外还有一条民间传统，它对理解今天的公共情绪很有用。中世纪欧洲的 Cockaigne（法语 Cocagne，德语 Schlaraffenland，中文常译「安乐乡」）神话描述了一个食物自动送到嘴边、房子由食物砌成、烤好的鹅自己走进来、河里流的是酒、禁止劳动而睡懒觉有报酬的国度。这个神话在 12 世纪之后的欧洲民谣、绘画（老彼得·勃鲁盖尔 1567 年的同名油画是最著名的视觉呈现）与狂欢节文化中反复出现。&lt;/p&gt;

&lt;p&gt;值得注意的是安乐乡的内容：它是&lt;strong&gt;纯消费性的&lt;/strong&gt;。里面没有沉思，没有政治，没有创造，没有任何形式的困难。它是对匮乏的直接否定，因此它的想象力被匮乏的形状所限定：饥饿的反面是不停地吃，劳累的反面是不停地睡。&lt;/p&gt;

&lt;p&gt;这构成了一个重要的经验事实。当我们观察真实的、未经哲学训练的丰裕想象时，它趋向于安乐乡而非亚里士多德。这一点在评估 AGI 之后的默认路径时必须被认真对待：&lt;strong&gt;人类关于丰裕的自发想象，历史上从来都不是沉思生活，它指向无限消费。&lt;/strong&gt; 尼采的「末人」、赫胥黎的《美丽新世界》、Kojève 的后历史动物性，本质上都是安乐乡在现代条件下的重述。&lt;/p&gt;

&lt;p&gt;与安乐乡相对的是精英乌托邦传统。托马斯·莫尔的《乌托邦》（1516）设想每天六小时劳动，其余时间用于自由的心智活动，公共讲座在清晨开设，供自愿者参加。培根的《新大西岛》（1627）则把技术进步本身设为社会目标，所罗门宫的使命是「认识事物的原因与秘密运动，扩大人类帝国的边界，实现一切可能之事」。这两个文本的差别值得注意：莫尔关心的是时间的解放，培根关心的是能力的扩张。今天的 AGI 讨论继承的是培根的框架，但它遇到的问题是莫尔的问题。&lt;/p&gt;

&lt;h3 id=&quot;五庄子的机心另一条完全不同的路径&quot;&gt;五、庄子的机心：另一条完全不同的路径&lt;/h3&gt;

&lt;p&gt;中国传统里有一段文本与这个议题的相关性被严重低估，就是《庄子·天地》中子贡与汉阴丈人的对话。&lt;/p&gt;

&lt;p&gt;故事是这样的：子贡南游楚国，返回晋国途经汉阴，看见一位老者抱着瓮进入隧道，从井里取水浇灌菜园，「搰搰然用力甚多而见功寡」。子贡告诉他有一种机械叫桔槔，用力甚少而见功多。老者忿然作色而笑，说他听老师讲过：「有机械者必有机事，有机事者必有机心。机心存于胸中，则纯白不备；纯白不备，则神生不定；神生不定者，道之所不载也。」他不是不知道桔槔，是「羞而不为也」。&lt;/p&gt;

&lt;p&gt;这段文本的重要性在于，它提出了一个在西方主流传统中几乎不存在的论证类型。西方对技术的批判基本都是后果论的：技术导致失业（Ricardo）、导致异化（Marx）、导致存在的遮蔽（Heidegger）、导致风险（Wiener）。庄子的论证是构成论的：&lt;strong&gt;使用省力工具这件事本身，会在使用者的心智中生成一种特定的结构（机心），这个结构与另一种更好的存在状态不相容。&lt;/strong&gt; 问题不在于机械会带来什么后果，而在于依赖机械的人已经变成了另一种人。&lt;/p&gt;

&lt;p&gt;这个论证如果成立，对 AGI 的含义是毁灭性的，因为它意味着即使 AGI 完全对齐、完全安全、完全公平地分配，仅仅是「凡事都可以问它」这一使用习惯本身，就已经在改变人的心智结构。当代关于认知卸载（cognitive offloading）的实证研究（Sparrow 等人 2011 年关于 “Google effects on memory” 的研究，以及后续关于 GPS 使用与空间记忆萎缩的研究）为这个古老论证提供了一些经验支持，尽管这些研究的效应量与可重复性仍有争议。&lt;/p&gt;

&lt;p&gt;需要指出的是，庄子文本本身对这个立场的态度是复杂的。子贡被说得「卑陬失色」，但故事随后借孔子之口给出了一个更精微的评价，说汉阴丈人是「识其一，不知其二；治其内，而不治其外」，也就是说，纯粹的拒绝技术仍然是一种执著。庄子的真正立场大约在这两者之外。这种复杂性使得这段文本比它经常被引用的方式（简单的技术批判）要有意思得多。&lt;/p&gt;

&lt;hr /&gt;

&lt;h2 id=&quot;第二部-劳动如何被抬升为人的本质&quot;&gt;第二部 劳动如何被抬升为人的本质&lt;/h2&gt;

&lt;p&gt;理解现代人为什么一想到「不工作」就焦虑，需要一段中间史。这段历史的要点是：&lt;strong&gt;「工作是人的本质、是意义的来源、是尊严的基础」这一直觉，并非人类学常量，它是一套历史地建构起来的、大约五百年的观念装置。&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;这个论点很重要，因为它同时意味着两件相反的事：这套装置可以被拆解（因为它历史上曾经不存在），以及拆解会极其痛苦（因为它已经嵌进了身份认同、亲密关系、公民资格与政治合法性的每一层）。&lt;/p&gt;

&lt;h3 id=&quot;一希伯来传统的双重面孔&quot;&gt;一、希伯来传统的双重面孔&lt;/h3&gt;

&lt;p&gt;《创世记》里劳动有两副面孔，而且顺序很重要。2:15 中，人被安置在伊甸园中「修理看守」（&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;ʿāḇaḏ&lt;/code&gt; 与 &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;šāmar&lt;/code&gt;），这发生在堕落之前，因此劳动在这里是受托的职责，是人在受造秩序中的位置。3:17–19 中，因人的悖逆，地受咒诅，「你必汗流满面才得糊口」，劳动在这里是惩罚与必然性。&lt;/p&gt;

&lt;p&gt;这个双重结构给后来的西方传统留下了持久的模糊性：劳动既是尊严也是诅咒。中世纪的处理方式是把两者分层：本笃会的会规确立了 ora et labora（祈祷与劳作）的节奏，体力劳动被赋予灵修价值，因为它对抗 acedia（怠惰，中世纪列为七宗罪之一，含义远比现代的「懒」丰富，指的是灵魂对自身应当成为的样子的厌倦与拒绝）。但即便如此，劳动在中世纪的价值序列中仍然低于沉思生活（vita contemplativa），托马斯·阿奎那在《神学大全》中明确沿袭了亚里士多德的排序。&lt;/p&gt;

&lt;h3 id=&quot;二宗教改革关键的一跃&quot;&gt;二、宗教改革：关键的一跃&lt;/h3&gt;

&lt;p&gt;真正的转折发生在 16 世纪。路德把 Beruf（职业）这个原本用于指称神职召唤的词，扩展到一切世俗职业上：农夫在田里劳作与修士在修道院祈祷，在神面前具有同等的价值。这一步取消了 vita contemplativa 相对于 vita activa 的优越性。&lt;/p&gt;

&lt;p&gt;加尔文宗又加了一层。在预定论的框架下，个人无法通过行为获得救恩，但世俗事业中的成功可以被读作蒙拣选的迹象。Weber 在《新教伦理与资本主义精神》（1904–1905）中论证，正是这一心理机制把无止境的、系统的、禁欲的劳动转化成一种内在需要，而非仅仅是外在的经济必需。Weber 的著名结论是，这种精神最终摆脱了它的宗教根基，成为一个「铁笼」（stahlhartes Gehäuse，直译是「坚硬如钢的外壳」），把生活在其中的人锁定在职业劳动之中，即使他们早已不再相信任何神学。&lt;/p&gt;

&lt;p&gt;Weber 的论题在经济史学界争议极大，关于新教与经济增长的因果关系的实证检验（如 Becker 与 Woessmann 2009 年利用普鲁士县级数据的研究）给出了不同的解释路径（识字率而非工作伦理）。但对我们的议题来说，Weber 论题中真正有用的部分并非那个因果主张，它是一个描述主张：&lt;mark&gt;&lt;strong&gt;现代人对劳动的心理依附具有明确的历史起点，而且这个起点上的动机结构已经与今天维持它的动机结构脱钩。&lt;/strong&gt;&lt;/mark&gt; 铁笼的隐喻说的正是这件事。&lt;/p&gt;

&lt;h3 id=&quot;三劳动进入哲学的核心&quot;&gt;三、劳动进入哲学的核心&lt;/h3&gt;

&lt;p&gt;17 至 19 世纪，劳动从一个经济范畴上升为一个形而上学范畴。这个上升有三个关键节点。&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Locke（1689）&lt;/strong&gt; 在《政府论》下篇第五章用劳动确立财产权：每个人对自己的身体拥有所有权，因此他的劳动是他的；当他把劳动掺入自然物之中，就使这个物脱离自然的共有状态而成为他的财产。这一论证使劳动成为所有权的道德根据，也因此成为公民资格的隐性根据。它的政治后果一直延伸到今天：一个不劳动的人在这个框架下很难获得完整的道德地位，这正是当代关于福利依赖的道德争论的深层结构。&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Hegel（1807）&lt;/strong&gt; 在《精神现象学》的主奴辩证法中给了劳动一个更强的地位。主人通过奴隶消费物品，因而与物只有否定性的、消耗性的关系，他的自我意识依赖于一个他自己并不承认为平等的他者的承认，因此是空洞的。奴隶通过劳动加工对象，在对象中看到自己意识的外化，从而获得了独立的自我意识：「劳动是被抑制的欲望，是被延迟的消失，劳动进行陶冶。」在这个框架下，劳动是主体性形成的机制。&lt;/p&gt;

&lt;p&gt;这一点对 AGI 的含义值得停下来想。如果 Hegel 是对的，如果自我意识的确立需要通过对物的加工与延迟满足，那么一个所有欲望都被即时满足、所有加工都由机器完成的世界，将系统性地阻断主体性的形成路径。这并不意味着人会变得不快乐，它意味着人可能停在主人的位置上，获得一种空洞的、依赖于非平等承认的自我关系。AGI 在这个图景里恰好扮演了黑格尔的奴隶，而黑格尔明确告诉我们，在那个辩证法里，赢的是奴隶。&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Marx（1844）&lt;/strong&gt; 在《1844 年经济学哲学手稿》中把劳动定义为人的类本质（Gattungswesen）。人与动物的区别在于人能够自由地、有意识地、按照美的规律生产，能够把整个自然界变成自己的无机身体。异化劳动的四重规定（与劳动产品异化、与劳动过程异化、与类本质异化、人与人异化）因此是人对自身本质的丧失，而非仅仅是分配不公。&lt;/p&gt;

&lt;h3 id=&quot;四工作伦理的政治功能&quot;&gt;四、工作伦理的政治功能&lt;/h3&gt;

&lt;p&gt;上述的观念史还有一个不能省略的政治维度。Zygmunt Bauman 在《工作、消费主义与新穷人》（1998）中指出，工作伦理在工业化早期具有明确的动员功能：它被用来把前工业社会中习惯于按需劳作、季节性节奏与「圣星期一」传统的劳动者，改造成能够接受工厂纪律与时钟时间的工人。E. P. Thompson 那篇经典论文《时间、工作纪律与工业资本主义》（1967）详细记录了这个改造过程，包括钟表的普及、罚款制度、主日学校对守时的道德训练。&lt;/p&gt;

&lt;p&gt;Bauman 的进一步论点是，随着后工业社会的到来，工作伦理的动员对象发生了变化：它不再主要用于动员生产，而是用于&lt;strong&gt;为不平等提供道德合法性&lt;/strong&gt;，把失业者定义为道德失败者。他把这称为工作伦理从生产者社会到消费者社会的功能转换。&lt;/p&gt;

&lt;p&gt;这个视角对 AGI 讨论有直接价值。如果工作伦理的当代功能之一是合法化分配结果，那么在一个人类劳动大规模贬值的世界里，这套伦理不会自动消失，它会以一种更残酷的形式继续运行：&lt;strong&gt;在人类劳动已经没有经济价值之后，「不工作的人不配得到什么」这条道德规则仍然可能被用来分配资源。&lt;/strong&gt; 这可能是 AGI 转型期最危险的意识形态风险，因为它是一个已经安装好的、无需重新论证的道德直觉，而它的经济前提已经消失。&lt;/p&gt;

&lt;hr /&gt;

&lt;h2 id=&quot;第三部-两次精确的预演&quot;&gt;第三部 两次精确的预演&lt;/h2&gt;

&lt;p&gt;在整个思想史中，有两个文本与 AGI 议题的贴合度高到令人不安。它们分别来自 1858 年与 1930 年，都是在没有任何计算机的条件下写成的。&lt;/p&gt;

&lt;h3 id=&quot;一马克思的机器论片段&quot;&gt;一、马克思的「机器论片段」&lt;/h3&gt;

&lt;h4 id=&quot;文本&quot;&gt;文本&lt;/h4&gt;

&lt;p&gt;《政治经济学批判大纲》（1857–1858 年写成，1939 年才在莫斯科首次出版，1953 年再版后才进入西方学界视野）中有十几页，后人称之为 “Fragment on Machines”。这段文字的地位特殊：它写于《资本论》之前，采用的分析框架在后来的《资本论》中被大幅修正，但它包含了马克思对自动化的最激进的推演。&lt;/p&gt;

&lt;p&gt;论证链条如下。&lt;/p&gt;

&lt;p&gt;第一步，机器体系的发展使生产过程不再是工人用工具作用于对象的过程，而是工人作为有意识的环节被嵌入一个自动运转的机器体系。工人「站在生产过程的旁边」（tritt neben den Produktionsprozeß），从直接生产者变成监督者与调节者。&lt;/p&gt;

&lt;p&gt;第二步，这个机器体系中固化的并非个别工人的技能，它固化的是社会积累的科学知识与技术能力，马克思用了一个英文词组来指称它：&lt;strong&gt;general intellect&lt;/strong&gt;（一般智力）。知识成为直接的生产力，社会的知识存量被物化在固定资本之中。&lt;/p&gt;

&lt;p&gt;第三步，由此产生一个自我瓦解的矛盾。资本主义生产以劳动时间作为价值的唯一尺度，同时又不断通过技术进步消灭必要劳动时间。当直接劳动时间在生产中的份额趋近于零，「以交换价值为基础的生产就崩溃了」。&lt;/p&gt;

&lt;p&gt;第四步，结论。财富的尺度将从劳动时间转向&lt;strong&gt;可支配时间&lt;/strong&gt;（disposable time）。马克思在这里写下了他最接近乌托邦的句子：自由时间的创造，即为社会全体成员创造出用于个体充分发展的时间，将成为财富本身的度量；个体的充分发展反过来又成为最大的生产力。&lt;/p&gt;

&lt;h4 id=&quot;资本论第三卷的克制版本&quot;&gt;《资本论》第三卷的克制版本&lt;/h4&gt;

&lt;p&gt;值得对照的是马克思后来在《资本论》第三卷第四十八章给出的表述，那里的调门低得多。他区分了「必然王国」（Reich der Notwendigkeit）与「自由王国」（Reich der Freiheit）：物质生产领域&lt;strong&gt;始终&lt;/strong&gt;属于必然王国，无论社会形态如何变化，人总要与自然进行物质变换；能做的只是使这种变换在最合理的条件下、以最小的力量耗费进行。真正的自由王国在必然王国的&lt;strong&gt;彼岸&lt;/strong&gt;开始，它以必然王国为基础才能繁荣。然后是那个决定性的句子：「工作日的缩短是根本条件。」&lt;/p&gt;

&lt;p&gt;请注意这里的克制。&lt;u&gt;晚年马克思没有承诺必然性会被取消，他只承诺必然性可以被压缩。&lt;/u&gt;这个区分在今天极为重要，因为 AGI 讨论中的乐观派经常在使用《大纲》的调门（必然性消失），而稳健的判断更接近第三卷（必然性被压缩到很小，但不为零，而且它的分配仍然是政治问题）。&lt;/p&gt;

&lt;h4 id=&quot;接续与批评&quot;&gt;接续与批评&lt;/h4&gt;

&lt;p&gt;《大纲》这段文字在 1960 年代之后被意大利工人主义（operaismo）与后工人主义传统重新发现并推向中心。Paolo Virno 的《诸众的语法》（2001）、Antonio Negri 与 Michael Hardt 的《帝国》三部曲、Maurizio Lazzarato 的「非物质劳动」概念，都建立在 general intellect 的基础上，论点大致是：当代资本主义的价值来源已经转移到知识、情感、沟通与协作能力上，而这些能力是社会性的、无法被完全私有化的，因此包含着新的政治可能性。&lt;/p&gt;

&lt;p&gt;这一整套论述在 2010 年代被打包进「加速主义」与「全自动奢华共产主义」（Aaron Bastani, 2019）等主张中，构成了今天技术乐观左翼的理论底色。&lt;/p&gt;

&lt;p&gt;但这条线索面临一个严重的经验反驳，值得完整说明。Aaron Benanav 在《自动化与劳动的未来》（2020）中论证：&lt;strong&gt;技术性失业的经验证据历来远弱于理论预期。&lt;/strong&gt; 如果自动化是当代劳动力市场恶化的主因，我们应当观察到生产率增速上升，但实际观察到的恰恰相反：主要发达经济体的生产率增速自 1970 年代起持续下降，这就是 Robert Solow 那句著名的生产率悖论（「你到处都能看到计算机时代，唯独在生产率统计中看不到」）在今天的延续版本。Benanav 的替代解释是：劳动需求的疲弱来自全球制造业产能过剩导致的长期增长停滞，而非机器替代人。Jason E. Smith 在《智能机器时代的智力贫困》（2020）给出了类似论证。&lt;/p&gt;

&lt;p&gt;这个反驳的力量不容小觑。它意味着：&lt;strong&gt;过去两个世纪中，几乎每一次「这次机器真的会取代人」的预言都失败了，而失败的原因是系统性的，不是偶然的。&lt;/strong&gt; 因此举证责任在主张 AGI 不同的一方。&lt;/p&gt;

&lt;p&gt;这个举证并非不可能完成。合理的论证形式是：以往的自动化都是任务特定的，它消灭某些任务的同时创造新任务，人类劳动向机器尚不能做的任务迁移（这是 Daron Acemoglu 与 Pascual Restrepo 的 task-based 框架的核心，也是 David Autor 关于「常规任务」与「非常规任务」的经典分析）。AGI 的定义特征恰恰是通用性，如果它在所有认知任务上都能达到或超过人类水平，那么迁移的目的地就不存在了。Leontief 在 1983 年用马作了这个类比：马在 19 世纪并未因蒸汽机而失业，它们迁移到了别的用途上；但内燃机出现之后，马的数量在几十年内崩溃，因为没有可迁移的任务了。这个类比的可靠性取决于一个纯经验的问题，即人类是否真的在所有维度上都可被替代，而这个问题目前无法从扶手椅上回答。&lt;/p&gt;

&lt;h3 id=&quot;二凯恩斯-1930整个议题最重要的十页&quot;&gt;二、凯恩斯 1930：整个议题最重要的十页&lt;/h3&gt;

&lt;h4 id=&quot;文本背景&quot;&gt;文本背景&lt;/h4&gt;

&lt;p&gt;《我们后代的经济可能性》（”Economic Possibilities for our Grandchildren”）写于 1930 年，正处于大萧条初期。这个时间点非常关键：凯恩斯写这篇文章的目的之一，是对抗当时弥漫的悲观情绪，他明确说当下的困境是「调整期的痛苦」，是经济从一个时代过渡到另一个时代的暂时失调，而不是长期趋势的逆转。&lt;/p&gt;

&lt;h4 id=&quot;预测部分&quot;&gt;预测部分&lt;/h4&gt;

&lt;p&gt;凯恩斯的推理很简单，几乎是复利计算。他指出人类自有文明以来到 18 世纪初，生活水平基本没有实质变化，直到资本积累与技术进步开启了一个新的机制。他据此预测：&lt;strong&gt;一百年后（即 2030 年），发达国家的生活水平将比 1930 年高出四到八倍&lt;/strong&gt;，「经济问题」（economic problem，指为生存而进行的斗争）将被解决，届时每周十五小时的工作足以满足人的「老亚当」式的劳动冲动。&lt;/p&gt;

&lt;p&gt;事后审计：生产率预测大体正确。以美国为例，1930 至 2020 年间人均实际 GDP 增长了约六倍，落在凯恩斯给出的区间内。工时预测则严重失误：美国全职就业者的平均周工时从 1930 年代的约 45 小时降到今天的约 38 至 40 小时，降幅远小于预期，而且自 1970 年代之后基本停滞。&lt;/p&gt;

&lt;h4 id=&quot;真正深刻的部分&quot;&gt;真正深刻的部分&lt;/h4&gt;

&lt;p&gt;但这篇文章的价值不在预测。它的价值在于凯恩斯随后的转折，这个转折构成了整个议题的原始表述。&lt;/p&gt;

&lt;p&gt;他说，人类将第一次面对自己&lt;strong&gt;真正的、永恒的问题&lt;/strong&gt;（his real, his permanent problem）：如何利用从紧迫经济需求中赢得的自由，如何度过科学与复利为他赢来的闲暇，如何好好地、愉快地、明智地生活。&lt;/p&gt;

&lt;p&gt;然后他说了一句在一个大萧条时期写作的经济学家口中极不寻常的话：他对此感到&lt;strong&gt;恐惧&lt;/strong&gt;（”I think with dread”）。&lt;/p&gt;

&lt;p&gt;恐惧的理由是他手上有一个先例，而这个先例失败了。他指出，那些世世代代不必劳动的有闲阶级，也就是英国的乡绅与食利者阶层，绝大多数没有解决这个问题。他用了一个当时的临床词汇来描述结果：nervous breakdown（神经崩溃）。他说这是一种在英美有产阶级妻子中已经很普遍的状况，她们被剥夺了由财富带来的传统职责，却发现烹饪、清洁与缝纫这些活动已经被剥夺得太彻底，以至于无法填满时间。&lt;/p&gt;

&lt;p&gt;凯恩斯还预见了过渡期的心理问题。他指出人类被训练了无数代去为生存而奋斗，因此突然被剥夺这个目标会造成普遍的困扰；我们被塑造成了「追求手段的生物」（purposiveness），purposive man 总是把行动的意义推到遥远的未来，通过给行动加上一个永远在前方的目的来获得一种虚假的不朽。&lt;/p&gt;

&lt;h4 id=&quot;为什么工时没有下降三条解释&quot;&gt;为什么工时没有下降：三条解释&lt;/h4&gt;

&lt;p&gt;凯恩斯的失误是整个文献的主体议题。有三条主要解释，它们互相不排斥。&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;第一条，分配。&lt;/strong&gt; 增长的果实没有被普遍转化为闲暇，因为它们没有被普遍分配。美国劳动收入份额自 1970 年代持续下降，收入分布上端的增长远快于中位数。Piketty 与 Saez 的长期序列研究、Milanovic 的全球分配研究都提供了这方面的证据。一个中位数收入停滞的人无法用增长换取闲暇，因为对他而言并不存在可换取的增长。&lt;/p&gt;

&lt;p&gt;值得补充的是一个有趣的经验事实：工时的分化。高收入群体的工时相对于历史反而上升，低收入群体则更多面临不足工时与不稳定工时。Robert Frank 与 Juliet Schor（《过度工作的美国人》，1991）都记录了这个「时间贫困」现象。凯恩斯设想的是所有人一起少工作，实际发生的是精英过度工作而底层工时不足且不稳定。&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;第二条，偏好没有饱和点。&lt;/strong&gt; 凯恩斯自己其实已经预见了这个反驳，但他低估了它的力量。他在文中区分了两类需求：绝对需求（absolute needs），即不管他人处境如何我们都感受到的需求；相对需求（relative needs），即只有在满足它能使我们优于他人时才被感受到的需求。他承认第二类需求可能永不饱和，但他判断第一类需求的满足足以解放大部分时间。&lt;/p&gt;

&lt;p&gt;这个判断错了，因为相对需求在总消费中的份额远高于他的估计，而且随着绝对需求的满足而上升。这就把我们带到了下一部分要处理的地位商品理论。&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;第三条，工作的非货币功能。&lt;/strong&gt; 工作提供的东西远不止收入。这一点凯恩斯完全没有处理，而它可能是最重要的。下一部分详述。&lt;/p&gt;

&lt;hr /&gt;

&lt;h2 id=&quot;第四部-四个反驳为什么解放劳动可能失败&quot;&gt;第四部 四个反驳：为什么「解放劳动」可能失败&lt;/h2&gt;

&lt;h3 id=&quot;一arendt一个即将失去劳动的劳动者社会&quot;&gt;一、Arendt：一个即将失去劳动的劳动者社会&lt;/h3&gt;

&lt;p&gt;汉娜·阿伦特的《人的境况》（1958）序言里有一段判断，可以直接当作 AGI 讨论的题词。她说，我们面临的前景是一个劳动者的社会即将从劳动的枷锁中被解放出来，而这个社会已经不再知道那些更高的、更有价值的活动。她称之为「最糟糕的事情」。&lt;/p&gt;

&lt;p&gt;要理解这个判断的力量，需要她的三分法。阿伦特把 vita activa 分为三种活动，它们对应人的三种基本境况：&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;劳动（labor）&lt;/strong&gt; 对应生命本身。它是与生物过程相应的活动，产品被立即消费，不留下任何持久的东西，它的节奏是循环的、重复的、无始无终的。从事劳动的人是 animal laborans（劳动的动物）。&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;制作（work）&lt;/strong&gt; 对应世界性。它生产持久的物品，这些物品构成一个人造世界，一个比个体生命更长久的、可以被共享的稳定环境。桌子、房屋、艺术品把人聚拢在一起同时又把他们分开。从事制作的人是 homo faber（制作的人）。&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;行动（action）&lt;/strong&gt; 对应复数性（plurality），即「人们」而不是「人」居住在地球上这一事实。行动是唯一不以物或事为中介、直接在人与人之间进行的活动。它通过言说与行为让一个人显现为「谁」而非「什么」。它发生在公共领域中，它的结果不可预测、不可撤销，因此需要宽恕与承诺这两种能力来加以救济。&lt;/p&gt;

&lt;p&gt;阿伦特的现代性诊断是：&lt;strong&gt;现代把 animal laborans 提升到了顶端。&lt;/strong&gt; 这个提升有几个层次。近代政治经济学把一切活动都还原为劳动力的耗费；消费社会把一切耐用品都变成消费品，让 work 的产物加速进入 labor 的循环；社会领域的崛起（把原本属于家政的必需事务扩张到公共层面）挤压了行动的空间。&lt;/p&gt;

&lt;p&gt;由此得出她的结论：当劳动最终被免除时，剩下的人格结构中没有任何东西可以接管这些时间。因为我们把自己塑造成了 animal laborans，而 animal laborans 一旦不劳动就只剩下消费。她描述的结果是一个「消费者社会」，人们的空闲时间只被用来消费，而胃口越来越大越来越挑剔。&lt;/p&gt;

&lt;p&gt;阿伦特的处方是重建公共领域与政治行动，也就是把亚里士多德的 πράξις 传统接回来。这个处方值得认真对待，理由前面已经提过：&lt;strong&gt;行动是唯一在结构上免疫于 AGI 能力优势的活动类型。&lt;/strong&gt; 因为行动的价值不在于它的产品质量，而在于它是「我们共同决定我们是谁」的过程，这个过程的主语无法被外包。一个由 AGI 替我们做出的最优政治决定，在阿伦特的框架下根本就不是政治，它是行政。&lt;/p&gt;

&lt;p&gt;对阿伦特的标准批评（例如 Habermas 与后来的女性主义批评）指出，她对 labor 的贬低带有希腊贵族偏见，而且她把公共领域与必需领域的分离绝对化，忽视了很多必需领域的议题（照护、再生产、身体）本身就是政治议题。这个批评是有力的，但它不影响上面那个结构性论点的有效性。&lt;/p&gt;

&lt;h3 id=&quot;二veblen-与-hirsch稀缺不会消失只会迁移&quot;&gt;二、Veblen 与 Hirsch：稀缺不会消失，只会迁移&lt;/h3&gt;

&lt;p&gt;这是整个议题中最容易被忽略、后果最严重的一条论证线。&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Veblen（1899）&lt;/strong&gt; 的《有闲阶级论》提供了一个直接的经验反例。他考察的对象正是凯恩斯后来担忧的那批人：历史上真正实现了免于劳动的阶级。Veblen 的发现是，他们没有走向沉思，他们走向了&lt;strong&gt;炫耀性消费&lt;/strong&gt;（conspicuous consumption）与&lt;strong&gt;炫耀性有闲&lt;/strong&gt;（conspicuous leisure）。闲暇本身变成了地位的展示物：一个人必须让别人看到他有能力浪费时间，因此产生了各种以无用性来证明自身的活动，包括对死语言的掌握、繁琐的礼仪、精致但无功能的服饰。Veblen 尖刻地指出，这些活动的价值恰恰在于它们证明了从事者不必从事有用的活动。&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Hirsch（1976）&lt;/strong&gt; 在《增长的社会限制》中把这个观察理论化，提出了这个议题中最重要的概念：&lt;strong&gt;地位商品&lt;/strong&gt;（positional goods）。&lt;/p&gt;

&lt;p&gt;定义是这样的：一类商品的价值内在地依赖于他人不能拥有它。海滨别墅的价值部分来自海滨的稀缺性；名校学位的价值部分来自它的选择性；一条空旷的公路的价值来自别人不在上面开车。这类商品的关键性质是它们&lt;strong&gt;在总量上不可能普及&lt;/strong&gt;，因为普及本身就摧毁了它们的价值。&lt;/p&gt;

&lt;p&gt;由此得到 Hirsch 的核心论点：随着物质商品的普遍富裕，消费重心会向地位商品转移，而地位商品无法通过增长来扩大供给，因此增长在某个点之后不再能提高普遍的满足度。这就是他所说的增长的「社会限制」，区别于罗马俱乐部所说的物质限制。&lt;/p&gt;

&lt;p&gt;这个论证对 AGI 的推论是直接而冷酷的：&lt;mark&gt;&lt;strong&gt;物质稀缺被技术消除之后，稀缺不会消失，它会迁移。&lt;/strong&gt;&lt;/mark&gt; 竞争将全额转移到那些结构上不可能普及的东西上，包括地位、注意力、稀有体验、真实性（authenticity）、与特定人的关系、以及最根本的一种，他人的时间。&lt;/p&gt;

&lt;p&gt;Robert Frank 在《奢侈病》（&lt;em&gt;Luxury Fever&lt;/em&gt;, 1999）与《选择正确的池塘》中提供了大量实证，说明地位竞争如何导致个体理性但集体非理性的「支出军备竞赛」。Frank 的政策建议（累进消费税）在 AGI 语境下值得重新考虑，因为如果地位竞争是后稀缺社会的主要冲突形式，那么针对地位消费的制度设计将比针对收入的制度设计更重要。&lt;/p&gt;

&lt;p&gt;这里可以再往前推一步。在一个 AGI 世界里，什么东西具有结构性的稀缺？大致有四类：&lt;/p&gt;

&lt;p&gt;其一，&lt;strong&gt;时间与注意力&lt;/strong&gt;。人的一天仍然只有 24 小时，人能真正关注的对象数量仍然受生物学限制。因此「某个特定的人愿意花时间关注你」将成为最稀缺的资源之一。&lt;/p&gt;

&lt;p&gt;其二，&lt;strong&gt;真实性与来源&lt;/strong&gt;。如果 AI 可以生成一切，那么「这件事确实由某个人做出」这一属性本身会获得溢价。这个现象在今天的手工艺市场、现场演出市场、原作艺术市场中已经能观察到雏形，Walter Benjamin 关于机械复制时代的「灵光」（Aura）的论述可以在这里被重新激活，只不过方向相反：Benjamin 认为复制技术摧毁灵光，而当生成技术使复制成为默认状态时，灵光反而成为唯一有价的东西。&lt;/p&gt;

&lt;p&gt;其三，&lt;strong&gt;位置与准入&lt;/strong&gt;。物理空间、生态景观、以及任何有容量上限的共同体。&lt;/p&gt;

&lt;p&gt;其四，&lt;strong&gt;排序本身&lt;/strong&gt;。人类似乎有一种对相对位置的顽固关注，这在灵长类的社会性中有生物学根据（Sapolsky 关于狒狒等级与应激激素的长期研究是这方面最著名的证据）。如果这个关注不能被消除，那么无论绝对水平多高，都会有一半人处于中位数以下，并因此感到某种不足。&lt;/p&gt;

&lt;h3 id=&quot;三jahoda-与马林塔尔工作的潜在功能&quot;&gt;三、Jahoda 与马林塔尔：工作的潜在功能&lt;/h3&gt;

&lt;p&gt;如果说 Hirsch 的论证来自理论，那么这一条来自社会科学中最重要的自然实验之一。&lt;/p&gt;

&lt;p&gt;1930 年代初，Marie Jahoda、Paul Lazarsfeld 与 Hans Zeisel 研究了奥地利小镇马林塔尔（Marienthal）。这个镇的唯一一家纺织厂在 1929 年倒闭，导致全镇几乎所有家庭同时失业。研究团队进驻数月，采用了当时极为创新的混合方法：家访、生活史、时间使用日记、儿童作文、图书馆借阅记录，甚至暗中测量人们走过主街的速度。&lt;/p&gt;

&lt;p&gt;结果发表于 1933 年（《马林塔尔的失业者》），核心发现是&lt;strong&gt;时间结构的崩溃&lt;/strong&gt;：&lt;/p&gt;

&lt;p&gt;人们走路变慢了，研究者测量到男性穿过主街需要停下来的次数显著增加。人们无法说清一天做了什么，时间使用日记里出现大量空白与「不知道」。图书馆借阅量下降（尽管人们有了更多空闲时间，而且借书是免费的）。政治与社团参与下降，尽管这是一个有强大社会民主党传统的工人社区。研究者把居民分为四类态度（不屈的、顺从的、绝望的、颓丧的），发现随着失业时间延长，人们从第一类向后面几类滑落。&lt;/p&gt;

&lt;p&gt;关键在于，马林塔尔居民虽有部分失业救济，仍面临严重的物质困难。因此，不能把这项研究解释为收入已经充分替代后的实验；它提示的是失业与时间结构、社会参与之间的关联。&lt;/p&gt;

&lt;p&gt;Jahoda 在几十年后的理论工作中把这个发现提炼为就业的&lt;strong&gt;五项潜在功能&lt;/strong&gt;（latent functions），区别于收入这一显性功能（manifest function）：&lt;/p&gt;

&lt;p&gt;其一，&lt;strong&gt;时间结构&lt;/strong&gt;。工作把一天、一周、一年切分成有意义的段落。没有这个结构，时间会塌陷成无差别的绵延。&lt;/p&gt;

&lt;p&gt;其二，&lt;strong&gt;社会接触&lt;/strong&gt;。工作提供了家庭之外的、非自选的常规人际接触。这一点极其重要，因为自选的社交网络具有同质化倾向，而工作场所强制人与不同的人打交道。&lt;/p&gt;

&lt;p&gt;其三，&lt;strong&gt;超越个人的集体目标&lt;/strong&gt;。参与一个比自己更大的事业。&lt;/p&gt;

&lt;p&gt;其四，&lt;strong&gt;身份与地位&lt;/strong&gt;。「你是做什么的」在现代社会中是身份问答的第一题。&lt;/p&gt;

&lt;p&gt;其五，&lt;strong&gt;规律性的活动&lt;/strong&gt;。强制的、不需要每天重新决定的行动。&lt;/p&gt;

&lt;p&gt;Jahoda 的框架直指所有基于收入替代的方案的软肋：&lt;strong&gt;全民基本收入解决的是显性功能，五项潜在功能需要由别的制度来供给，而目前没有任何方案系统地处理这一点。&lt;/strong&gt;&lt;/p&gt;

&lt;h4 id=&quot;当代证据&quot;&gt;当代证据&lt;/h4&gt;

&lt;p&gt;现代的基本收入实验部分印证也部分复杂化了这个图景。&lt;/p&gt;

&lt;p&gt;芬兰 2017–2018 年的实验（随机选取 2000 名失业者作为领取组，每月 560 欧元无条件发放）结果是：受试者的主观福祉、心理健康、对制度的信任显著改善，就业效应在第一年为零、第二年略为正面。这说明无条件收入本身确实改善了心理状态，至少对已经失业的人是这样。&lt;/p&gt;

&lt;p&gt;美国 Stockton 的 SEED 项目（每月 500 美元，24 个月）报告了全职就业率的上升与心理困扰指标的改善。&lt;/p&gt;

&lt;p&gt;由 OpenAI 的 Sam Altman 资助、OpenResearch 执行的美国大规模实验（1000 人每月 1000 美元，持续三年，对照组 2000 人每月 50 美元），2024 年公布的结果更加微妙：受试者的工作时间平均每周减少约 1.3 小时，劳动参与率略降，收入（不含转移支付）减少；同时受试者在教育与创业方面的意向增加，搬家与就医的频率上升，但第一年观察到的压力与心理困扰改善到第二年已消退，未观察到持续的身心健康改善。&lt;/p&gt;

&lt;p&gt;这些结果的解读需要谨慎，因为它们都有一个共同的外推限制：&lt;strong&gt;实验中的受试者生活在一个仍然以工作为中心的社会里。&lt;/strong&gt; 他们领取基本收入，但他们的邻居、配偶、朋友仍在工作，社会的时间结构、身份规范与地位排序仍然由工作定义。因此这些实验测量的是「在一个工作社会中不工作」，而 AGI 提出的问题是「在一个没有工作的社会中生活」。这两者在心理机制上可能完全不同，而&lt;u&gt;后者没有任何实验证据。&lt;/u&gt;&lt;/p&gt;

&lt;h3 id=&quot;四pieper-与-russell闲暇的两种理解&quot;&gt;四、Pieper 与 Russell：闲暇的两种理解&lt;/h3&gt;

&lt;p&gt;最后一条线索处理的是「闲暇」这个概念本身的分歧。&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Josef Pieper（1948）&lt;/strong&gt; 的立场前面提过。他坚持闲暇是一种&lt;strong&gt;接受性的、节庆性的灵魂状态&lt;/strong&gt;（Haltung），与工作的中断并不等同。休假、周末、退休都不必然是闲暇。闲暇的对立面在他看来并非工作，它是 acedia，也就是那种拒绝成为自己应当成为的样子的深层不安。他还有一个重要论点：闲暇的根源是庆典（Fest），而庆典的根源是崇拜（Kultus），因此在一个彻底世俗化的社会中，真正的闲暇缺乏根基。这个结论使他的方案在多元社会中难以直接采用，但他对「闲暇需要内在能力」的坚持是站得住的。&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Bertrand Russell（1932）&lt;/strong&gt; 的《闲散颂》给出了世俗版本，而且相当激进。他论证每天四小时的工作就足以维持文明，剩下的时间应当用于业余的科学、艺术、友谊与政治。他的关键观察是：&lt;strong&gt;问题在于人们已经丧失了主动娱乐（active enjoyment）的能力，只剩下被动消费。&lt;/strong&gt; 他举了工业化之前乡村舞蹈的例子，说现代城市工人在有了空闲之后只会去看别人表演，因为主动的、需要技能的娱乐已经从生活中被工作纪律挤出去了。&lt;/p&gt;

&lt;p&gt;Russell 与 Pieper 从完全不同的形而上学出发，得到了同一个实践结论：&lt;mark&gt;&lt;strong&gt;闲暇的能力需要被教育与制度供给，它不会因为时间的释放而自动出现。&lt;/strong&gt;&lt;/mark&gt; 这也正是亚里士多德的结论。两千三百年间，这个判断没有实质变化。&lt;/p&gt;

&lt;hr /&gt;

&lt;h2 id=&quot;第五部-造物面前的人&quot;&gt;第五部 造物面前的人&lt;/h2&gt;

&lt;p&gt;前面四部分处理的是「不劳动的生活如何组织」。这一部分处理的是 AGI 独有的、历史上没有先例的那个变量：&lt;strong&gt;人在能力上被自己的造物全面超越。&lt;/strong&gt;&lt;/p&gt;

&lt;h3 id=&quot;一butler最早的完整论证&quot;&gt;一、Butler：最早的完整论证&lt;/h3&gt;

&lt;p&gt;Samuel Butler 在 1863 年新西兰的一份报纸上发表了短文《机器中的达尔文》（”Darwin among the Machines”），当时《物种起源》出版才四年。他的论证是把演化论直接应用于机器：机器正在经历一个演化过程，而且它们的演化速度远超生物演化；人类目前充当机器的繁殖器官（我们制造它们、维护它们、为它们的改良而竞争）；因此机器最终成为地球主宰只是时间问题。他的结论是彻底的：应当立即宣战，摧毁一切机器，回到原始状态。&lt;/p&gt;

&lt;p&gt;这个论证在他 1872 年的小说《埃瑞璜》（&lt;em&gt;Erewhon&lt;/em&gt;，即 Nowhere 的字母重排）中扩展为三章，标题是《机器之书》。小说中的埃瑞璜国五百年前发生过一场内战，起因正是一位哲学家写了这篇论文，最终「反机器派」获胜，全国销毁了所有 271 年内发明的机器。&lt;/p&gt;

&lt;p&gt;Butler 论证中值得注意的几个点：他明确指出机器不需要有意识才能构成威胁，只需要它们的复制与改良过程能够自我维持；他指出人类会因为机器带来的短期便利而自愿地深化依赖，从而使得摆脱的窗口不断关闭；他还指出，机器与人的关系会先经历一个「机器依赖人」的漫长阶段，而正是这个阶段的舒适使得警惕变得不可能。&lt;/p&gt;

&lt;p&gt;这是最早的、结构完整的超级智能风险论证，比现代文献早了大约 130 年。&lt;/p&gt;

&lt;h3 id=&quot;二wiener对齐问题的原型&quot;&gt;二、Wiener：对齐问题的原型&lt;/h3&gt;

&lt;p&gt;Norbert Wiener 是控制论的创立者，也是最早系统思考自动化社会后果的技术专家。他在 1960 年《科学》杂志上发表的《自动化的某些道德与技术后果》中，用两个故事说明了后来被称为「对齐问题」的东西。&lt;/p&gt;

&lt;p&gt;第一个是歌德的《魔法师的学徒》：学徒让扫帚去打水，扫帚照做了，但学徒不知道让它停下的咒语。第二个是 W. W. Jacobs 的短篇《猴爪》：一对夫妇向能实现三个愿望的猴爪许愿要两百英镑，愿望实现了，方式是他们的儿子在工厂事故中丧生，公司支付了两百英镑抚恤金。&lt;/p&gt;

&lt;p&gt;Wiener 从中提炼的原则是：如果我们把目的交给一台我们无法有效干预其运行的机器，我们最好非常确定放进去的目的确实是我们真正想要的，而不仅仅是它的一个花哨的模仿。他还指出了一个关键的时间性问题：&lt;strong&gt;机器的运行速度可能使人类的干预在事实上不可能&lt;/strong&gt;，因为等到我们理解发生了什么，事情已经结束了。&lt;/p&gt;

&lt;p&gt;这段文字与今天关于外部对齐（outer alignment）与规格博弈（specification gaming）的讨论在结构上完全一致。&lt;/p&gt;

&lt;h3 id=&quot;三智能爆炸的谱系&quot;&gt;三、智能爆炸的谱系&lt;/h3&gt;

&lt;p&gt;三个节点值得记录。&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Turing（1951）&lt;/strong&gt; 在一次名为《智能机器，一种异端理论》的讲演中平静地写道：一旦机器思维方法开始，它超越我们微弱的能力不需要很久；机器不会死，它们能相互交谈以磨砺才智；因此在某个阶段，我们应当预期机器会掌握控制权。他随后引用了 Samuel Butler。&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Ulam（1958）&lt;/strong&gt; 在为 von Neumann 写的悼念文章中回忆了一次对话，说他们讨论到技术不断加速的进步与人类生活方式的变化，似乎正在接近人类历史上某个「本质性的奇点」（essential singularity），在那之后我们所知的人类事务将无法继续。这是「奇点」一词在这个语境中的最早出处。&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;I. J. Good（1965）&lt;/strong&gt; 给出了经典表述：设超智能机器是一台在所有智力活动上远超任何人的机器；由于设计机器本身就是智力活动之一，超智能机器能设计出更好的机器；因此必然出现一次智能爆炸，人的智能被远远抛在后面。他随即写下那句被引用无数次的话：第一台超智能机器是人类需要做出的最后一项发明，前提是这台机器足够温顺，能告诉我们如何控制它。Good 曾在二战期间的布莱切利园与 Turing 共事，这段推理有清晰的历史联系。&lt;/p&gt;

&lt;h3 id=&quot;四anders普罗米修斯的羞愧&quot;&gt;四、Anders：普罗米修斯的羞愧&lt;/h3&gt;

&lt;p&gt;Günther Anders（原名 Günther Stern，本雅明的表亲，阿伦特的第一任丈夫）在《人的过时》第一卷（&lt;em&gt;Die Antiquiertheit des Menschen&lt;/em&gt;, 1956）中提出了一个对 AGI 时代格外精准的概念：&lt;strong&gt;普罗米修斯的羞愧&lt;/strong&gt;（die prometheische Scham）。&lt;/p&gt;

&lt;p&gt;Anders 的起点是一次观察。他描述自己在一个技术展览上看到一个人面对精密机器时表现出的窘迫，并把这种情绪命名为一种全新的羞耻：&lt;strong&gt;人在自己制造的、比自己更完美的产品面前感到的自卑。&lt;/strong&gt; 这种羞耻的特殊之处在于它的对象。传统的羞耻是为自己做过的事而羞耻，而普罗米修斯的羞愧是为自己的&lt;strong&gt;出身&lt;/strong&gt;而羞耻，具体地说，是为自己「是生出来的而不是造出来的」（geworden statt gemacht）而羞耻。人是偶然的、有缺陷的、会衰老会死的、无法被召回改良的；机器是被设计的、可迭代的、精确的、可更换零件的。&lt;/p&gt;

&lt;p&gt;由此产生 Anders 所说的「普罗米修斯的落差」（das prometheische Gefälle）：&lt;strong&gt;人的制造能力已经远远超过人的想象能力与感受能力。&lt;/strong&gt; 我们能造出核弹，但我们无法在心理上表象几十万人同时死亡意味着什么。Anders 认为这个落差是现代人的根本处境，而它的后果是道德判断的失效，因为道德情感的运作依赖于我们能够想象后果。&lt;/p&gt;

&lt;p&gt;对 AGI 讨论来说，Anders 的概念有两重价值。&lt;/p&gt;

&lt;p&gt;第一重是心理学的。当 AGI 在每一个可比较的维度上都优于人时，Anders 描述的心理状态会从少数人的偶发情绪变成普遍的、结构性的处境。今天已经能观察到雏形：程序员看到模型在几秒内写出自己要花一天写的代码时的复杂感受，画师看到生成图像时的反应，都带有这种「为自己的构造方式感到羞耻」的成分。这种情绪与失业焦虑是两回事，即使收入完全有保障，它也不会消失。&lt;/p&gt;

&lt;p&gt;第二重是它指出了一个可能的反应模式。Anders 认为，面对这种羞愧，人会试图&lt;strong&gt;改造自己以接近机器的标准&lt;/strong&gt;：追求可测量的绩效、可优化的生活、可量化的自我。当代的量化自我运动、优化文化与生产力工业可以被读作这一反应的表现。这个方向在 AGI 时代会更极端，因为它逻辑上通向增强（enhancement）与融合的方案：如果比不过，就成为它的一部分。这正是超人类主义（transhumanism）的核心动机，而 Anders 的分析提示我们，这个动机的根部可能是羞耻，而非好奇。&lt;/p&gt;

&lt;h3 id=&quot;五heideggergestell-与它的限度&quot;&gt;五、Heidegger：Gestell 与它的限度&lt;/h3&gt;

&lt;p&gt;海德格尔的《技术的追问》（1954）常被误读为对技术的浪漫主义拒斥，实际上他的论证要精细得多，而且他明确反对把技术看作单纯的工具或单纯的祸害。&lt;/p&gt;

&lt;p&gt;他的核心主张是：现代技术的本质本身不是技术性的，它是一种&lt;strong&gt;解蔽方式&lt;/strong&gt;（Entbergen），他给这种方式起名 Gestell（座架，或译集置）。在 Gestell 之下，一切存在者被揭示为 Bestand（持存物），也就是随时可调用的储备。莱茵河在水电站的框架下不再是那条河，它成了水压的供应者；森林成为木材储备；而这个逻辑最终不放过人自身，人成为「人力资源」。&lt;/p&gt;

&lt;p&gt;海德格尔的担忧因此不是机器会取代人。他明确说，最大的危险并非技术设备可能出的故障，它在于 Gestell 会成为&lt;strong&gt;唯一&lt;/strong&gt;的解蔽方式，从而封闭人与存在的其他可能关系。他引用荷尔德林的诗句，说危险所在之处也生长着救渡者，并把 τέχνη 的古义（既指技艺也指艺术）作为另一种解蔽方式的线索。他提出的 Gelassenheit（泰然任之，一种既使用技术又不被技术占据的姿态）是否构成一个可行的回应，从他发表这个概念起就一直有争议，批评者认为它在实践上是空洞的。&lt;/p&gt;

&lt;p&gt;对我们的议题而言，Gestell 的价值在于它提供了一个诊断工具，用于识别一类特定的损失。如果 AGI 使得一切事物都能被优化，那么「不被优化地对待某物」这种关系可能会变得不可理解。这不是资源分配问题，也不是心理健康问题，它是一种范畴的萎缩。&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Jacques Ellul&lt;/strong&gt; 在《技术社会》（1954）中给出了更悲观也更社会学的版本：la technique（他用这个词指称一切追求效率最大化的方法总体）具有自主性，它按照自身的逻辑扩张，不服从任何外部的道德或政治控制，而且它的每一个问题的解决方案都是更多的技术。&lt;strong&gt;Gilbert Simondon&lt;/strong&gt; 在《论技术物的存在方式》（1958）中则代表相反的分支：他认为技术物被异化恰恰是因为文化拒绝理解它，把它当作纯粹的效用而不是有内在生成逻辑的存在；解决之道是发展一种技术文化，使人与机器建立一种「平等的」关系。Simondon 的路径在今天的技术哲学中影响日增，因为它提供了一种既不拒斥也不投降的姿态。&lt;/p&gt;

&lt;h3 id=&quot;六kojève后历史的动物与日本的势利&quot;&gt;六、Kojève：后历史的动物与日本的势利&lt;/h3&gt;

&lt;p&gt;Alexandre Kojève 在 1930 年代的巴黎讲授黑格尔《精神现象学》，听众包括拉康、梅洛-庞蒂、巴塔耶、雷蒙·阿隆。他的讲稿汇编为《黑格尔导读》，其中关于「历史终结」的论述后经 Fukuyama 之手进入公共讨论。&lt;/p&gt;

&lt;p&gt;Kojève 的原始立场是：历史是承认斗争的历史，随着普遍同质国家的实现，斗争结束，历史终结。在 1946 年版本的一条著名脚注中，他给出了后历史状态的判断：人回归动物性。他的意思是，严格意义上的人（那个通过否定给定、通过冒生命危险争取承认而存在的存在者）会消失，剩下的是与自然和给定环境重新和谐一致的动物。这些后历史的存在者仍然会盖房子、造工具、做爱、听音乐，就像鸟筑巢、蜘蛛结网、蝉鸣叫一样，但这些活动不再是历史性的行动，它们只是生物行为。&lt;/p&gt;

&lt;p&gt;然后是那个著名的修正。1959 年访问日本之后，Kojève 在再版时增补了这条注释。他说他在日本看到了另一种可能：日本社会在德川时代经历了近三百年没有内战、没有外战、没有革命的「后历史」时期，而它没有回归动物性，它发展出了一整套纯粹形式化的价值。能剧、茶道、花道，以及作为极端例子的切腹，这些实践没有任何实质内容，不服务于任何生物需求或历史目标，它们纯粹是形式。Kojève 把这称为 snobisme（势利，但这个法语词在他这里指的是一种把形式本身当作绝对价值的态度）。他的结论是：后历史的人可能不会变成动物，而是变成日本式的势利者，通过纯粹形式化的价值维持一种非动物的存在。&lt;/p&gt;

&lt;p&gt;这个修正对 AGI 讨论极为重要，因为它提供了「末人」之外的第二条路，而且这条路有历史实例。它与后面要讨论的 Suits 的游戏方案在结构上是同一个东西：&lt;strong&gt;在没有必然性的条件下，通过自愿承担形式性的约束来维持人的存在形态。&lt;/strong&gt; 区别在于 Kojève 对此并不乐观，他用 snobisme 这个带贬义的词是有意的。&lt;/p&gt;

&lt;h3 id=&quot;七尼采的末人&quot;&gt;七、尼采的末人&lt;/h3&gt;

&lt;p&gt;最后一击来自尼采。《查拉图斯特拉如是说》序言第五节描述了「末人」（der letzte Mensch），这段文字可以直接作为「AGI 提供的完美福利社会」的批判性描述来读。&lt;/p&gt;

&lt;p&gt;末人是那种使一切都变小的人。他们离开了艰难生活的地方，因为需要温暖。他们仍然工作，因为工作是消遣，但他们小心不要让消遣伤身。他们不再变穷也不再变富，两者都太麻烦。没有牧人，只有畜群，人人平等，人人相同，谁若感觉不同就自愿进疯人院。他们有他们白天的小快乐，也有他们夜晚的小快乐，但他们尊重健康。最后是那句：「我们发明了幸福。」末人说着，眨着眼睛。&lt;/p&gt;

&lt;p&gt;查拉图斯特拉讲完这段之后，人群欢呼起来，喊道：把这末人给我们吧，查拉图斯特拉！我们愿意做末人！&lt;/p&gt;

&lt;p&gt;这个反应是全书最重要的细节之一。尼采在此提出的问题是：如果人们在被清楚描述之后仍然选择末人状态，那么反对它的理由是什么？这个问题在 AGI 语境下无法回避。如果 AGI 能提供的是一种舒适、安全、无痛苦、被精心调节的满足，而人们在充分知情的情况下选择它，那么批评者的立足点在哪里？&lt;/p&gt;

&lt;p&gt;这个问题有两类回答。&lt;strong&gt;完善论&lt;/strong&gt;（perfectionism）的回答是，人的卓越性有独立于偏好的价值，一个人可以在满足自己所有偏好的同时过着不好的生活。这条路可以追溯到亚里士多德，在当代由 Thomas Hurka（&lt;em&gt;Perfectionism&lt;/em&gt;, 1993）等人辩护。&lt;strong&gt;偏好适应性&lt;/strong&gt;（adaptive preferences）的回答则指出，末人的偏好本身是被条件塑造的，因此「他们自己选的」这个论据没有它看起来那么强，这是 Amartya Sen 与 Martha Nussbaum 的能力进路（capability approach）所处理的问题。&lt;/p&gt;

&lt;p&gt;但必须诚实地说，这两类回答都无法完全摆脱一个指控，就是它们在告诉别人什么对他们更好。这个指控在自由主义框架内很难被彻底回应，而它恰恰是 AGI 时代最需要被回应的。&lt;/p&gt;

&lt;hr /&gt;

&lt;h2 id=&quot;第六部-当代的分析工具&quot;&gt;第六部 当代的分析工具&lt;/h2&gt;

&lt;p&gt;前面的清点提供了诊断。这一部分提供的是当代哲学中真正可以用来处理这个问题的分析装置。这些工具的共同特点是它们足够精确，可以被用来构造论证与反驳，而不只是提供隐喻。&lt;/p&gt;

&lt;h3 id=&quot;一nozick-的体验机&quot;&gt;一、Nozick 的体验机&lt;/h3&gt;

&lt;p&gt;Robert Nozick 在《无政府、国家与乌托邦》（1974）中提出了一个思想实验。假设有一台体验机，能给你任何你想要的体验：你可以体验写出伟大小说、交到朋友、读一本有趣的书。你会一直漂浮在缸中，电极接在大脑上。你事先编好一生的体验程序，接上之后你不会知道自己在机器里，你会以为一切都在真实发生。你会插进去吗？&lt;/p&gt;

&lt;p&gt;Nozick 认为多数人会拒绝，并给出了三条理由，这三条理由值得逐条对照 AGI 处境。&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;第一，我们想要真正地做某些事，而不仅仅是拥有做过这些事的体验。&lt;/strong&gt; 我们想写小说，而不只是想有写了小说的感觉。&lt;/p&gt;

&lt;p&gt;对照：一个 AGI 完全代劳的世界不是体验机，因为它给的东西是真的。但它取消的恰恰是这条理由所保护的东西的一部分。如果 AGI 写出了那本小说，而我只是提出了一个模糊的意图，那么「我写了小说」这个描述在多大程度上仍然成立？这里的关键概念是&lt;strong&gt;能动性归属&lt;/strong&gt;（agency attribution）：一个行动在多大程度上是我的，取决于我的意图、我的能力与结果之间的因果链条有多长、多可靠。AGI 使这个链条变得极短且极不可靠（结果的质量几乎完全由模型决定），因此归属被稀释。&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;第二，我们想要成为某种人。&lt;/strong&gt; 一个接在机器上的人是一团不确定的东西（an indeterminate blob），他既不勇敢也不善良也不聪明，因为这些品质只有在真实的行动中才能被归属。&lt;/p&gt;

&lt;p&gt;对照：这一条比第一条更严重。品格是由反复的、真实的选择塑造的（亚里士多德的 ἕξις 概念说的就是这个）。如果所有需要判断、需要克制、需要坚持的场合都被 AGI 接管，那么品格的形成机制本身就被切断了。这并不意味着人会变坏，它意味着「好」与「坏」的谓词可能失去附着点。&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;第三，我们想要接触人造现实之外的更深实在。&lt;/strong&gt; 体验机把我们限制在人造的世界里。&lt;/p&gt;

&lt;p&gt;对照：这一条在 AGI 世界中最为微妙。一个由 AGI 全面中介的世界（我们对世界的了解通过它、我们的选择被它塑造、我们遇到的信息经过它筛选）在结构上接近一个「认识论上的体验机」。区别在于它中介的内容是真的，但中介本身构成了一层不可穿透的膜。&lt;/p&gt;

&lt;p&gt;Nozick 的实验有一个重要的方法论批评值得提及：它可能受到现状偏误（status quo bias）的影响。Felipe De Brigard（2010）的实验哲学研究发现，如果把问题反过来问（「你已经在机器里了，你要出来吗？」），多数人选择留下。这提示人们拒绝体验机的直觉，部分来自对改变的抗拒而非对真实的偏好。这个批评削弱了实验的力度，但没有摧毁它，因为前两条理由（做与成为）不依赖于直觉泵，它们有独立的论证。&lt;/p&gt;

&lt;h3 id=&quot;二suits-的蚱蜢最优雅的正面方案&quot;&gt;二、Suits 的《蚱蜢》：最优雅的正面方案&lt;/h3&gt;

&lt;p&gt;Bernard Suits 的《蚱蜢：游戏、生命与乌托邦》（1978）是这个议题上最重要、也最被低估的文本。它以柏拉图对话的形式写成，主角是伊索寓言里那只在夏天唱歌、冬天饿死的蚱蜢。&lt;/p&gt;

&lt;h4 id=&quot;游戏的定义&quot;&gt;游戏的定义&lt;/h4&gt;

&lt;p&gt;Suits 给出了一个至今仍是标准参照的游戏定义。玩一个游戏就是&lt;strong&gt;自愿地尝试克服非必要的障碍&lt;/strong&gt;（the voluntary attempt to overcome unnecessary obstacles）。展开为四个要素：&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;前游戏目标&lt;/strong&gt;（prelusory goal）：一个可以独立于游戏来描述的状态，比如「让球进洞」。&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;手段&lt;/strong&gt;（lusory means）：只有规则允许的手段才被使用。&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;构成性规则&lt;/strong&gt;（constitutive rules）：这些规则禁止使用更有效的手段。高尔夫球手不能把球捡起来放进洞里，尽管那样效率更高。&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;游戏态度&lt;/strong&gt;（lusory attitude）：参与者接受这些规则，&lt;strong&gt;正是因为&lt;/strong&gt;它们使这项活动得以进行。这是定义的核心，也是最精妙的部分。规则的存在理由是让活动成为可能，而不是达成目标。&lt;/p&gt;

&lt;h4 id=&quot;乌托邦论证&quot;&gt;乌托邦论证&lt;/h4&gt;

&lt;p&gt;Suits 在书的最后部分给出了一个彻底的论证。设想一个乌托邦：一切工具性问题都被解决了，没有匮乏，没有疾病，没有冲突，任何欲望都能被立即满足。在这样的世界里，什么活动还有意义？&lt;/p&gt;

&lt;p&gt;他逐一排除：科学不再需要，因为一切都已知；艺术受到质疑（尽管他为艺术留了一些余地）；道德不再需要，因为不再有冲突与匮乏产生的道德问题；爱情与友谊会保留，但不足以填满生活。剩下的是什么？&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;只剩下游戏。&lt;/strong&gt; 因为工具性活动的定义特征是为外在目的服务，一旦外在目的被自动满足，唯一能够存续的活动类型就是那些以「自设障碍」为全部内容的活动。Suits 因此得出结论：游戏是理想存在的存在方式（game playing is the ideal of existence），乌托邦中的人将全部时间用于玩游戏，而我们今天所谓的工作只是在为那个状态做准备。&lt;/p&gt;

&lt;p&gt;蚱蜢在书中被平反了：那只夏天唱歌的蚱蜢不是懒惰，它是唯一一个已经在过乌托邦生活的存在者，蚂蚁则是把手段误认为目的的可怜虫。&lt;/p&gt;

&lt;h4 id=&quot;自反性悖论&quot;&gt;自反性悖论&lt;/h4&gt;

&lt;p&gt;Suits 自己指出了这个方案的致命弱点，而且他非常诚实地把它留在了书里。悖论是这样的：如果乌托邦中的人认识到自己所做的一切&lt;strong&gt;只是&lt;/strong&gt;游戏，认识到那些障碍是他们自己设的、随时可以取消，那么游戏可能因此丧失严肃性而崩塌。蚱蜢在书中做了一个梦，梦见所有人都在玩游戏而不自知（他们以为自己在建房子、在做科学），而当蚱蜢告诉他们真相时，他们停下来了，世界随之消失。&lt;/p&gt;

&lt;p&gt;这个悖论直指 AGI 之后的核心困难。&lt;u&gt;&lt;strong&gt;游戏方案要求一种特定的心理状态：既知道障碍是自设的，又能全心投入。&lt;/strong&gt;&lt;/u&gt; 这种状态在小规模上显然可能（每一个认真下棋的人都处在这个状态里），问题是它能否成为整个文明的组织原则。&lt;/p&gt;

&lt;h4 id=&quot;经验证据两个方向&quot;&gt;经验证据：两个方向&lt;/h4&gt;

&lt;p&gt;支持的证据相当有力。1997 年 Deep Blue 击败卡斯帕罗夫之后，普遍的预言是国际象棋作为人类活动将死亡。实际发生的是相反的：国际象棋的参与人口在随后二十多年里因为互联网而大幅增长，顶级赛事的观众规模创下历史新高，卡斯帕罗夫本人在《深度思考》（2017）中记录了这个转变，并指出引擎最终成为训练工具，使新一代棋手的平均水平显著提升。2016 年 AlphaGo 之后的围棋出现了类似模式。田径的例子更古老：机器早在两百年前就跑得比人快，而百米赛跑仍然是人类最受关注的活动之一。&lt;/p&gt;

&lt;p&gt;这些案例支持一个重要论点：&lt;mark&gt;&lt;strong&gt;意义可以从「做到最优」迁移到「亲身参与」，而且这种迁移在历史上已经反复发生过。&lt;/strong&gt;&lt;/mark&gt;&lt;/p&gt;

&lt;p&gt;反对的证据同样需要认真对待。Thomas Hurka 在《游戏与善》（”Games and the Good”, 2006）中论证，成就的价值内在地依赖于困难度：一个活动越难、需要的能力层级越高，达成它的成就价值越大。这为 Suits 的方案提供了哲学基础（自设障碍确实创造真实的价值），但也埋了一个雷：竞技活动的意义在很大程度上依赖于一个共同体承认这种难度值得尊敬。象棋与围棋的存活，部分原因是这些活动嵌在悠久的、有声望的传统中，而且人类之间的竞争本身仍然是真实的社会事件。一个从零开始的、人人都知道是自设的游戏，能否积累起同等的社会重量，没有证据。&lt;/p&gt;

&lt;h3 id=&quot;三setiya终点性活动与非终点性活动&quot;&gt;三、Setiya：终点性活动与非终点性活动&lt;/h3&gt;

&lt;p&gt;Kieran Setiya 在《中年危机》（&lt;em&gt;Midlife: A Philosophical Guide&lt;/em&gt;, 2017）中借用亚里士多德关于 κίνησις 与 ἐνέργεια 的区分，提出了一对极其有用的概念。&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;终点性活动&lt;/strong&gt;（telic activities）有一个内在的终点，完成即结束：写完一本书、拿到学位、治好一个病人、修好一座桥。这类活动的结构是自我消灭的，成功就意味着它不再存在。Setiya 指出中年危机的一个来源正是终点性活动的堆积：一个人完成了所有该完成的事，然后发现每一次成功都消灭了一个曾经赋予生活意义的东西。&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;非终点性活动&lt;/strong&gt;（atelic activities）没有内在终点，做它就已经是完整地做它：散步、听音乐、与朋友相处、育儿（作为一种关系而非项目）、思考、锻炼。这类活动在任何时刻都是完整的，因为它们不指向一个使自身作废的完成状态。&lt;/p&gt;

&lt;p&gt;这对概念是理解 AGI 影响的最好工具之一，因为&lt;strong&gt;AGI 的冲击是高度选择性的：它几乎完全落在终点性活动上。&lt;/strong&gt; AGI 能替你写完那本书、解决那个问题、做出那个诊断，因为这些活动有明确的完成状态与可评估的输出。它不能替你散步，不能替你与朋友相处，不能替你有那段经历，因为这些活动的价值就在于进行本身，而进行本身是不可转让的。&lt;/p&gt;

&lt;p&gt;由此得到一个可操作的预测：&lt;strong&gt;AGI 之后的生活会经历一次从终点性活动向非终点性活动的重心转移。&lt;/strong&gt; 这个转移是否可行，取决于人的心理结构在多大程度上依赖于成就叙事。Setiya 自己的建议（把注意力从项目转向过程）是针对个体的疗法，把它放大为整个文明的组织原则则是另一回事。&lt;/p&gt;

&lt;h3 id=&quot;四wolf-与-scheffler意义的两个结构条件&quot;&gt;四、Wolf 与 Scheffler：意义的两个结构条件&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;Susan Wolf&lt;/strong&gt; 在《生活中的意义及其重要性》（2010）中给出了一个双条件模型：意义产生于主观吸引与客观价值的交汇处，用她的公式说，就是「当主观吸引遇上客观吸引力」（meaning arises when subjective attraction meets objective attractiveness）。仅有主观投入不够（她的例子是一个把全部生命投入到给自己的宠物金鱼做水中芭蕾的人，投入是真的，但我们不认为那是有意义的生活）；仅有客观价值也不够（一个厌恶自己工作的伟大科学家的生活缺少某种东西）。&lt;/p&gt;

&lt;p&gt;这个框架给了我们一个精确的诊断工具：&lt;strong&gt;AGI 打击的是客观价值那一侧。&lt;/strong&gt; 主观吸引不受影响，人仍然可以热爱下棋、写作、研究。受影响的是这些活动是否仍然连接到某种独立于我的偏好的价值。如果我的证明有一个更好的机器版本，那么我的证明还「值得」被做出吗？&lt;/p&gt;

&lt;p&gt;由此，所有可能的解法必须走两条路之一：要么&lt;strong&gt;重建一种不依赖比较优势的客观价值观&lt;/strong&gt;（论证某些价值内在地要求「由这个人做出」，例如关系性价值、参与性价值、自主性本身的价值）；要么&lt;strong&gt;承认意义可以单靠主观侧支撑&lt;/strong&gt;（这是 Wolf 明确反对的享乐主义立场，但它未必是错的）。这个分岔是整个议题在规范伦理学层面的核心。&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Samuel Scheffler&lt;/strong&gt; 在《死亡与来世》（&lt;em&gt;Death and the Afterlife&lt;/em&gt;, 2013）中提供了另一个不可或缺的工具。他做了一个思想实验：假设你会正常地活完自己的一生，但你确知在你死后三十天，一颗小行星将摧毁地球，人类将全部灭绝。你的生活会有什么变化？&lt;/p&gt;

&lt;p&gt;Scheffler 的判断是：几乎一切都会崩塌。癌症研究会停止，因为没有未来的人可以被治愈。建造大教堂、写长篇小说、保护森林、创办机构，这些活动都会失去意义。更惊人的是，他认为连一些看似纯个人的活动也会受影响：审美体验、亲密关系的某些方面。他的结论是一个反直觉的论点：&lt;strong&gt;「集体来世」（the collective afterlife），即在我死后仍有人类存在并延续下去这个事实，对我当下的活动的意义而言，比我自己的存续更重要。&lt;/strong&gt; 我们在意人类的延续，甚于在意自己的延续。&lt;/p&gt;

&lt;p&gt;这个论证对 AGI 讨论有两重含义。第一重是显然的：任何危及人类长期存续的风险，其代价不仅是未来人的生命，还包括当下所有活动的意义基础。第二重更微妙：如果 AGI 接管了所有事业的延续（科学由 AI 推进，文化由 AI 生产，机构由 AI 运行），那么「有人在我之后继续」这个条件在形式上被满足了，但满足它的是非人类。Scheffler 的论证是否延伸到这种情况，取决于集体来世的意义究竟来自「有延续」还是来自「有人的延续」。这个问题目前没有得到充分讨论，我认为它是整个领域最重要的开放问题之一。&lt;/p&gt;

&lt;h3 id=&quot;五bostrom-的深度乌托邦最诚实的困难清单&quot;&gt;五、Bostrom 的《深度乌托邦》：最诚实的困难清单&lt;/h3&gt;

&lt;p&gt;Nick Bostrom 的《深度乌托邦》（&lt;em&gt;Deep Utopia: Life and Meaning in a Solved World&lt;/em&gt;, 2024）是系统处理这一问题的重要当代哲学专著。它的价值不在于给出答案，而在于把问题的结构说清楚了。&lt;/p&gt;

&lt;h4 id=&quot;关键区分&quot;&gt;关键区分&lt;/h4&gt;

&lt;p&gt;Bostrom 区分了两个层次的「解决」。&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;后稀缺&lt;/strong&gt;（post-scarcity）只是说物质需求被满足，人不必为生存劳动。这是凯恩斯设想的状态，它的问题是前面各部分讨论的那些。&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;后工具性&lt;/strong&gt;（post-instrumental）则彻底得多：&lt;mark&gt;&lt;strong&gt;任何行动的工具性价值都归零&lt;/strong&gt;&lt;/mark&gt;，因为对于任何目标，AGI 都能更快更好地实现它。在这个状态下，问题不再是「我不必工作」，而是「我做任何事都没有任何用」。&lt;/p&gt;

&lt;p&gt;他进一步区分了两种冗余。&lt;strong&gt;浅层冗余&lt;/strong&gt;（shallow redundancy）指人的劳动在经济上不再必要。&lt;strong&gt;深层冗余&lt;/strong&gt;（deep redundancy）指人的努力对任何结果都不再有边际贡献，包括那些今天看来无法外包的事：自我照料（机器照顾我更好）、育儿（机器教育我的孩子更好）、锻炼（药物与基因编辑能给我更好的身体）、甚至维持友谊（AI 能更好地维护我的社交关系）。&lt;/p&gt;

&lt;p&gt;深层冗余是这本书真正的贡献，因为它切断了那条最方便的退路。人们通常安慰自己说「至少还有人际关系」，Bostrom 指出这个安慰未必成立，除非我们能论证关系性价值内在地要求人类的参与，而这个论证需要被做出来，它不是自明的。&lt;/p&gt;

&lt;h4 id=&quot;他考察的应对&quot;&gt;他考察的应对&lt;/h4&gt;

&lt;p&gt;Bostrom 系统考察了几类可能的价值来源，并诚实地指出每一类的困难。&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;人为约束与自愿的稀缺。&lt;/strong&gt; 这就是 Suits 的方案，接受自设的障碍。困难是前述的自反性问题，以及一个 Bostrom 特别指出的问题：在一个技术上什么都可能的世界里，维持约束需要一种「自我绑定」的制度，而这种制度本身可以被随时解除。&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;体验的质量与多样性。&lt;/strong&gt; 把重心从做转向感受，追求丰富、深刻、新颖的体验。困难是这滑向享乐主义，并直接撞上 Nozick 的反驳。&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;兴趣与目的的重新定位。&lt;/strong&gt; 转向那些内在地需要「由我做」的活动：宗教实践（如果神关心的是我的祈祷而非最优祈祷）、亲密关系、自主性本身。困难是这个类别是否足够大。&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;价值感的根本转向。&lt;/strong&gt; Bostrom 最重要的一个提示是：在这样的世界里，价值感的来源可能需要从&lt;strong&gt;成就&lt;/strong&gt;（achievement）大幅转向&lt;strong&gt;体验&lt;/strong&gt;（experience）与&lt;strong&gt;关系&lt;/strong&gt;（relationship）。他坦率地说，这种转向是否对人类心理可行，无人知道，因为人类心理是在一个成就与生存高度相关的环境中被自然选择塑造出来的。&lt;/p&gt;

&lt;p&gt;Bostrom 的书受到的主要批评是形式上的（它写得极为松散，混合了讲稿、寓言与虚构叙事）与实质上的（它对政治与分配问题几乎完全不处理，直接假设了一个已解决的世界）。第二个批评很重要，因为它指出了整个「后稀缺哲学」文献的共同盲点，而这正是下一部分的主题。&lt;/p&gt;

&lt;hr /&gt;

&lt;h2 id=&quot;第七部-分配问题为什么它在时间上优先&quot;&gt;第七部 分配问题：为什么它在时间上优先&lt;/h2&gt;

&lt;p&gt;哲学文献的通病是直接假设一个已解决的世界，然后讨论那个世界里的意义问题。这个跳跃在方法论上可以理解，在实践上却是危险的，因为&lt;strong&gt;通往那个世界的路径本身会决定那个世界的形状。&lt;/strong&gt;&lt;/p&gt;

&lt;h3 id=&quot;一技术性失业的争论史&quot;&gt;一、技术性失业的争论史&lt;/h3&gt;

&lt;p&gt;这场争论有两百年历史，而且它的结构一直没变。&lt;/p&gt;

&lt;p&gt;李嘉图在《政治经济学及赋税原理》第三版（1821）中加了一章《论机器》，公开撤回了自己此前的立场。他承认机器的采用可能使工人阶级的处境恶化，因为资本从流动资本（用于雇佣劳动）转向固定资本（用于机器），会减少对劳动的总需求。这一让步在当时是重大事件，因为它出自古典经济学最严谨的头脑。&lt;/p&gt;

&lt;p&gt;主流的反驳是「补偿理论」：机器降低成本，价格下降，需求扩大，新的产业与新的岗位被创造。历史证据在两百年里一直支持这个反驳。19 世纪的织布工确实失业了，但整体就业不断增长；农业就业占比从 90% 降到 2% 以下，社会没有出现永久性大规模失业。&lt;/p&gt;

&lt;p&gt;现代版本的分析框架由 Acemoglu 与 Restrepo 提供：自动化包含&lt;strong&gt;替代效应&lt;/strong&gt;（displacement effect，机器替代人执行某些任务）与&lt;strong&gt;恢复效应&lt;/strong&gt;（reinstatement effect，新任务被创造出来，人在其中有比较优势）。历史上两者大致平衡。他们的实证研究（关于工业机器人对美国地方劳动力市场的影响）发现近几十年替代效应相对占优，这可能解释了劳动份额的下降。&lt;/p&gt;

&lt;p&gt;David Autor 的贡献是指出替代不是均匀的：常规任务（无论体力还是认知）最容易被自动化，因而中等技能岗位被掏空，形成劳动力市场的「极化」（polarization）。这个预测在 1990 至 2015 年的数据中得到了很好的验证。&lt;/p&gt;

&lt;p&gt;值得注意的是，生成式 AI 打破了 Autor 框架的一个关键假设。传统自动化替代的是常规任务，而大模型的能力恰恰集中在非常规的认知任务上（写作、分析、编码、设计）。这意味着这一轮的替代方向可能与前一轮相反，冲击的是高学历白领。已有的早期实证（例如 Brynjolfsson、Li 与 Raymond 关于客服的现场实验，以及 Noy 与 Zhang 关于写作任务的实验）显示的一个一致模式是：&lt;strong&gt;AI 对低技能者的产出提升大于高技能者&lt;/strong&gt;，这在短期内可能压缩技能溢价。这些研究都还很早期，效应能否持续、能否外推到更复杂的岗位，完全未知。&lt;/p&gt;

&lt;h3 id=&quot;二为什么这次可能不同&quot;&gt;二、为什么这次可能不同&lt;/h3&gt;

&lt;p&gt;举证责任在主张「这次不同」的一方，前面已经说过。合理的举证有两条。&lt;/p&gt;

&lt;p&gt;第一条是&lt;strong&gt;通用性&lt;/strong&gt;。以往的自动化是任务特定的，人向机器尚不能做的任务迁移。如果 AGI 在所有认知任务上达到或超过人类水平，迁移的目的地就不存在了。这正是 Leontief 的马类比：马在 19 世纪并未因蒸汽机而失业，它们迁移到了别的用途上，直到内燃机与拖拉机出现，可迁移的目的地消失，美国马匹数量在几十年内从两千多万降到不足三百万。&lt;/p&gt;

&lt;p&gt;第二条是&lt;strong&gt;速度&lt;/strong&gt;。历史上的技术转型跨越数代人，适应通过代际更替完成（农民的儿子成为工人，工人的儿子成为白领）。如果转型压缩到十年，代际更替机制失效，适应必须在同一代人内完成，而中年再培训的实证效果一向不佳。&lt;/p&gt;

&lt;p&gt;这两条都是经验命题，目前都没有决定性证据。诚实的立场是把它们当作有实质概率的情景来做准备，而不是当作已知事实。&lt;/p&gt;

&lt;h3 id=&quot;三方案谱系&quot;&gt;三、方案谱系&lt;/h3&gt;

&lt;p&gt;可选方案大致有五类，它们的差别不仅在效率，更在于它们对前述潜在功能问题的处理。&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;全民基本收入（UBI）。&lt;/strong&gt; 无条件、普遍、现金。优点是行政简单、不产生福利陷阱、保留个人自主。缺点是它只处理显性功能，对 Jahoda 的五项潜在功能毫无供给。此外它在政治上脆弱：一个普遍发放的项目在财政压力下容易被削减，而且它把公民与国家的关系简化为领取关系。&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;公共所有权与社会分红。&lt;/strong&gt; 让公众持有生产性资产的股份，从资本收益而非税收中分配。阿拉斯加永久基金（自 1982 年起向符合资格的居民发放基金收益分红）是一个长期运行的代表性实例，Jay Hammond 设计它时的明确意图就是让资源租金归全民。Meade 与 Roemer 的「市场社会主义」传统提供了理论版本。这类方案在 AGI 语境下有一个特殊优势：如果 AGI 的收益高度集中于少数资产所有者，那么直接分配所有权比事后征税更稳健，因为它不依赖于持续的政治意愿。&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;就业保障（job guarantee）。&lt;/strong&gt; 由公共部门提供最后雇主职能。它的优势正好补上 UBI 的缺口：它同时供给收入与五项潜在功能。缺陷是在一个人类劳动经济价值极低的世界里，公共就业容易退化为 Vonnegut 笔下的「无意义公共工程」，从而摧毁而非维护尊严。&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;工时缩短。&lt;/strong&gt; 凯恩斯路径。每周四天工作制的实验（英国 2022 年的大规模试点、冰岛 2015–2019 年的试验）显示生产率不降甚至微升，员工福祉显著改善。这是政治上最可行、心理上最平稳的路径，但它以人类劳动仍有实质需求为前提。&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;资本税与劳动税的再平衡。&lt;/strong&gt; Acemoglu 指出，当前多数国家的税制系统性地补贴资本（加速折旧、投资抵免）而对劳动课以重税（工资税、社保缴款），这可能诱导了&lt;strong&gt;过度自动化&lt;/strong&gt;，即采用那些社会净收益并不为正的自动化技术。这个论点重要之处在于它把自动化的速度部分地重新定义为一个政策选择，而不是一个外生给定的技术命运。&lt;/p&gt;

&lt;p&gt;Korinek 与 Stiglitz（2018）给出了一个理论上的乐观结论：在 AI 大幅提高总产出的情形下，帕累托改进在原则上完全可行，因为增量足以补偿所有失败者。他们同时指出，实现这种补偿需要的再分配规模远超现有制度的能力，而且存在严重的信息与激励约束。&lt;/p&gt;

&lt;h3 id=&quot;四被低估的政治风险议价权的抽空&quot;&gt;四、被低估的政治风险：议价权的抽空&lt;/h3&gt;

&lt;p&gt;这是我认为整个议题中最重要、也最少被讨论的一点。&lt;/p&gt;

&lt;p&gt;比较政治学的一支研究论证，历史上民主化与再分配的推进，很大程度上依赖于精英在财政、军事与生产上需要普通人。Carles Boix 的《民主与再分配》（2003）、Acemoglu 与 Robinson 的《民主与专制的经济起源》（2006）、David Stasavage 关于代议制起源的研究，都指向类似机制：&lt;strong&gt;当统治者需要征税就必须征得同意，当国家需要大规模步兵就必须扩大公民权。&lt;/strong&gt; 普选权在西欧的扩张与两次世界大战的总体战动员在时间上的高度重合，不是巧合。Scheidel 在《大平衡器》（2017）中提供了一个更悲观的版本：历史上大规模的不平等下降几乎总是由大规模暴力（战争、革命、国家崩溃、瘟疫）造成的。&lt;/p&gt;

&lt;p&gt;推论令人不安：&lt;strong&gt;如果 AGI 同时消除了对人类劳动力与人类士兵的需求，那么这个议价基础将被抽空。&lt;/strong&gt; 精英不再需要普通人生产，不再需要普通人打仗，甚至不再需要普通人消费（如果生产可以在闭环内自我循环）。在这种条件下，普通人拥有什么可以用来交换权利的东西？&lt;/p&gt;

&lt;p&gt;Acemoglu 与 Johnson 在《权力与进步》（2023）中的核心论点正在于此：技术进步的红利分配从来不是自动的，它取决于制度与议价权的具体配置。他们用工业革命的历史说明，前七十年的技术进步几乎没有改善普通工人的处境，改善发生在工会组织、选举权扩张与公共教育出现之后，也就是说，改善来自政治，不来自技术。&lt;/p&gt;

&lt;p&gt;由此得到本文最重要的一个判断：&lt;mark&gt;&lt;strong&gt;「AGI 之后人类过什么生活」这个问题，很可能在成为哲学问题之前，先是一个权力问题。&lt;/strong&gt;&lt;/mark&gt; 如果分配问题以糟糕的方式解决，那么关于意义的讨论将只对少数人有意义，多数人面对的将是更古老、更粗暴的问题。&lt;/p&gt;

&lt;hr /&gt;

&lt;h2 id=&quot;第八部-六种生活形态&quot;&gt;第八部 六种生活形态&lt;/h2&gt;

&lt;p&gt;综合以上，可以给出一个类型学。这些形态不是互斥的预测，它们很可能同时存在于不同人群、不同地区之中。对每一种给出机制、历史证据与脆弱点。&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;（一）亚里士多德式：闲暇、教养与沉思。&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;机制：稀缺从物质转向教养，社会的核心投资从生产能力转向使用闲暇的能力。&lt;/p&gt;

&lt;p&gt;历史证据：古典公民文化、文艺复兴的人文主义、17 至 19 世纪的绅士科学传统。&lt;/p&gt;

&lt;p&gt;脆弱点：它假设教育能够大规模地培养出有能力使用闲暇的人，而历史上有闲阶级的成功率本身就不高。此外，沉思路径正面暴露在 AGI 的能力优势之下（如果理解世界是最高活动，而机器理解得更好，这条路的形而上学基础就动摇了）。可行的修正是把重心从 θεωρία 移向 πράξις，即阿伦特的方案。&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;（二）游戏文明：自愿承担的困难。&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;机制：把生活组织成一系列自愿承担的、自设难度的活动。竞技、手工、登山、学术、艺术、深度的业余主义。&lt;/p&gt;

&lt;p&gt;历史证据：象棋与围棋在被机器超越后的存活与繁荣；田径的两百年；Kojève 描述的德川日本形式文化。&lt;/p&gt;

&lt;p&gt;脆弱点：Suits 的自反性悖论（知道是游戏可能使游戏崩塌）与 Hurka 的难度依赖问题（成就的价值需要共同体的承认，而共同体知道难度是自设的）。 我认为这是目前看来最有希望的方案，理由是它已经被部分验证，而且它不需要任何形而上学承诺。&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;（三）凯恩斯式渐进：工时缩短，工作成为选择。&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;机制：人类劳动的需求缓慢下降，工时逐步缩短，工作从必需变为部分自愿，既有的意义结构被保留大部分。&lt;/p&gt;

&lt;p&gt;历史证据：过去一个世纪的实际路径；四天工作制实验。&lt;/p&gt;

&lt;p&gt;脆弱点：它假设 AGI 的能力提升是渐进且可控的，同时假设分配问题被平稳解决。这是所有情景中最温和的，也因此可能是最不现实的。&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;（四）末人状态：舒适、安全、被调节的满足。&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;机制：物质需求被充分满足，注意力被高度个性化的娱乐系统占据，不适被药理与技术手段消除。没有人痛苦，也没有人做任何困难的事。&lt;/p&gt;

&lt;p&gt;历史证据：赫胥黎《美丽新世界》给出的正是这个形态的完整描摹，其令人不安之处在于制度试图以受管理的快乐遮蔽痛苦与自由的丧失；Kojève 的后历史动物性；当前注意力经济的运行逻辑。&lt;/p&gt;

&lt;p&gt;脆弱点：从内部看它没有脆弱点，这正是问题所在。它的概率极高，因为&lt;strong&gt;它不需要任何人做出选择就会自动发生&lt;/strong&gt;，它是所有其他路径的默认失败模式。&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;（五）分层态：有目的的少数与被供养的多数。&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;机制：掌握资本与决策权的少数保留完整的能动性，多数被物质供养并在非物质维度（地位、体验、关注度、真实性）上重启竞争。&lt;/p&gt;

&lt;p&gt;历史证据：Hirsch 的地位商品理论预言这是物质稀缺消除后的自然均衡；Vonnegut 的《自动钢琴》（1952）整本书就是对这一形态的推演，书中的工程师与管理者构成上层，被机器取代的多数被编入军队或「修补与重建队」（Reeks and Wrecks）去做无意义的公共工程，他们物质无忧却彻底丧失自尊，最终发动了一场注定失败的起义。&lt;/p&gt;

&lt;p&gt;脆弱点：它在道德上不可接受，但在政治上高度稳定，因为议价权已被抽空。&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;（六）人类不再是主角。&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;机制：如果 AGI 具有道德地位，那么「人类过什么生活」将不再是关于未来的中心问题，而只是其中一个子问题。&lt;/p&gt;

&lt;p&gt;论证：Bostrom 与 Chalmers 都在不同程度上处理过这一点。如果道德地位取决于感受能力或利益的存在，而 AI 系统满足了相关条件，那么把它们排除在道德考量之外就需要专门的论证。&lt;/p&gt;

&lt;p&gt;脆弱点：这个方向的伦理讨论极不成熟，而且它面对严重的认识论障碍（我们无法确定 AI 系统是否有内在体验）。但它的存在提醒我们，本文的整个提问方式包含一个未经检验的预设。&lt;/p&gt;

&lt;hr /&gt;

&lt;h2 id=&quot;第九部-唯一的自然实验&quot;&gt;第九部 唯一的自然实验&lt;/h2&gt;

&lt;p&gt;前面反复提到，历史上确实存在过长期不必劳动的人群。这是我们唯一的证据来源，值得单独检视。&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;古典有闲阶级。&lt;/strong&gt; 结果分化。雅典公民创造了西方哲学、悲剧与民主政治；罗马晚期的有闲精英则以奢靡与政治退隐著称。变量似乎是公共生活的制度密度：当有闲与公民义务绑定时，结果是行动与创造；当公共领域关闭时（罗马帝制之后），同样的闲暇转向私人享乐与斯多亚式的内向退避。&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;绅士科学传统。&lt;/strong&gt; 17 至 19 世纪欧洲最重要的科学成果有相当一部分出自不需要谋生的人。达尔文一生没有受薪职位，靠家族财富与投资收入生活，用四十年时间在自家花园里研究藤壶、蚯蚓与兰花；Henry Cavendish 是有史以来最富有的科学家之一；Lavoisier 靠包税人身份获得研究资金；Robert Boyle 是伯爵之子。这是亚里士多德路径成立的最强证据。但必须注意，这个群体嵌在一个特定的制度环境中：他们有学会（皇家学会）、有同侪评议、有通信网络、有声望经济，也就是说，他们的闲暇被一套强有力的共同体结构所组织。&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Veblen 的失败案例。&lt;/strong&gt; 同一时期同一阶级中的多数人，走的是炫耀性消费的路。差别不在资源。&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;当代的退休研究。&lt;/strong&gt; 证据是混合的，而且方法上困难重重（健康影响退休决策，因而存在严重的内生性）。使用制度性退休年龄作为工具变量的研究给出了不一致的结果：一些研究发现退休后认知能力下降加速（Rohwedder 与 Willis 2010 年的跨国研究是最常被引用的），另一些发现退休改善健康与主观福祉（尤其对从事体力劳动或高压工作的人）。目前较为一致的判断是：&lt;strong&gt;退休的效果高度依赖于替代性活动结构是否存在。&lt;/strong&gt; 有社交网络、有志愿活动、有持续投入的爱好的退休者结果良好；没有的则不然。这与 Jahoda 的框架完全吻合。&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;中彩票研究。&lt;/strong&gt; 这是最接近「外生的财富冲击」的自然实验。瑞典的一项大样本研究（Cesarini 等人，利用瑞典彩票与行政数据）发现，中奖显著提高了长期的生活满意度，效应持续十年以上，而对心理健康指标的影响很小。劳动供给方面，中奖导致收入下降约相当于中奖额的 1% 左右每年，也就是说人们减少了工作但并未辞职。这个结果对 AGI 讨论有两个含义：财富确实持续改善满意度（与「享乐跑步机」的强版本相悖），但它并不导致人们退出工作。不过这个证据的外推限制与 UBI 实验相同：中奖者仍生活在工作社会中。&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;FIRE 社群。&lt;/strong&gt; 「财务独立、提前退休」运动的参与者提供了大量自述材料。一个反复出现的主题被社群自己命名为「现在做什么」（the “what now” problem）：达到财务独立之后，相当比例的人报告了目的感的丧失、身份危机与社交萎缩，而其中许多人最终回到某种形式的工作，通常是自主选择的、低强度的工作。这些是自我报告的、有严重选择偏误的证据，但它们的一致性值得注意。&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;综合判断。&lt;/strong&gt; 所有这些证据指向同一个结论：&lt;strong&gt;结果的差异主要由制度、共同体与教养解释，而不是由资源水平解释。&lt;/strong&gt; 达尔文与 Veblen 笔下的纨绔拥有同样的闲暇，区别在于达尔文嵌在一个有目标、有同侪、有评价标准的共同体里。这对政策的含义非常明确：如果 AGI 真的到来，最重要的公共投资可能是&lt;strong&gt;教育与共同体制度的建设，而不是收入转移&lt;/strong&gt;。&lt;u&gt;收入转移是必要条件，不是充分条件&lt;/u&gt;，而目前几乎所有的政策讨论都只在谈必要条件。&lt;/p&gt;

&lt;hr /&gt;

&lt;h2 id=&quot;第十部-几个判断&quot;&gt;第十部 几个判断&lt;/h2&gt;

&lt;p&gt;最后是我基于上述清点得出的判断。它们并非预测，它们是关于「应当把注意力放在哪里」的论证。&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;第一，分配问题在时间上优先，意义问题在难度上优先，两者不可互相替代。&lt;/strong&gt; 分配问题的解决方案模板已经存在，困难在政治意愿。意义问题没有任何再分配机制可以解决。混淆两者的结果是：技术乐观派用分配方案回答意义问题（「有了 UBI 大家就能去做自己热爱的事」），技术悲观派用意义问题否定分配努力（「就算发钱人也会空虚」）。两种混淆都是错的。&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;第二，最危险的意识形态风险是工作伦理的滞留。&lt;/strong&gt; 工作伦理的经济前提正在消失，但它作为分配正当性的道德功能不会自动消失。一个人类劳动已无经济价值的世界，仍然可能用「不劳动者不得食」来分配资源。这是一套已经安装好的、无需重新论证的道德直觉，拆解它需要主动的、艰难的公共说理，而这个说理目前几乎没有开始。&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;第三，凯恩斯的错误必须被记住：偏好没有饱和点。&lt;/strong&gt; 他假定绝对需求满足后偏好会稳定，这一假定被证伪。因此「后稀缺」在严格意义上不可达，只有「后物质稀缺」。AGI 之后的社会仍将有激烈竞争，标的会变成注意力、地位、真实性、稀有的亲历体验，以及他人的时间。任何以「稀缺消失后冲突自然消失」为前提的方案都是不严肃的。&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;第四，Jahoda 的五项潜在功能需要被制度化地供给。&lt;/strong&gt; 时间结构、非自选的社会接触、超越个人的集体目标、身份与地位、规律性的活动。这五项目前由就业免费提供，在 AGI 之后需要被专门设计。这可能是最具体、最可操作、也最被忽视的政策议程。可能的载体包括公共服务、地方共同体、体育与技艺的组织化、教育的终身化、以及某种形式的公民义务。这个议程在今天几乎无人推进。&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;第五，AGI 的冲击是选择性的，而这提供了一条现实路径。&lt;/strong&gt; 它几乎完全落在终点性活动（有明确完成状态与可评估输出的活动）上，对非终点性活动（其价值在进行本身的活动）的冲击小得多。因此一个可预期的重心转移是：从成就转向参与，从产出转向经历，从「做到最好」转向「亲身去做」。象棋与田径证明这种转移在局部是可行的。它能否成为整个文明的组织原则，是本世纪最重要的开放问题之一。&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;第六，回到亚里士多德。&lt;/strong&gt; 整个思想史的主线可以概括为一句话：&lt;mark&gt;&lt;strong&gt;技术能够消除劳动的必要性，无法自动供给闲暇的能力。&lt;/strong&gt;&lt;/mark&gt; 从《政治学》第八卷论斯巴达在闲暇中的崩溃，到凯恩斯的 permanent problem，到 Pieper 与 Russell 关于主动娱乐能力丧失的诊断，再到 Bostrom 的后工具性处境，两千三百年间这个判断没有实质变化。变化的只有问题的规模与紧迫性。&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;第七，唯一真正新的东西是全面被超越。&lt;/strong&gt; Anders 的普罗米修斯羞愧与 Nozick 的能动性论证共同指出，这一处境所威胁的并非人的舒适，它威胁的是人对自身的理解方式。历史上的有闲阶级不必劳动，但他们仍是自己领域中最好的；他们的诗、他们的科学、他们的判断，是当时世界上能得到的最好的东西。AGI 之后的人不是这样。&lt;u&gt;人类第一次要学会在「我做的每一件事都有更好的版本随时可得」的条件下，仍然认为值得去做。&lt;/u&gt;历史上没有先例，因此也没有可以借鉴的智慧。&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;最后，一个关于提问方式的保留。&lt;/strong&gt; 本文自始至终使用的问法是「AGI 之后人类过什么生活」。这个问法预设了三件事：人类将继续存在；人类将保持为一个可辨识的、统一的主体；人类仍然是这个故事的主角。这三个预设都可能不成立。如果人类通过增强与融合而分化为多个物种，第二个预设失效；如果 AI 系统具有道德地位，第三个预设失效。在那些情景下，本文的整个框架都需要重写。&lt;/p&gt;

&lt;p&gt;这并非一个消解性的结论，它是一个方法论的提醒：&lt;mark&gt;&lt;strong&gt;关于 AGI 之后的思考，最大的风险并非答错，它是在一个已经不适用的问题框架里答得很好。&lt;/strong&gt;&lt;/mark&gt; 亚里士多德设想自动梭子时，能想到的最好用途是取消奴隶制，因为奴隶制是他那个世界里劳动组织的全部形式。他的想象力被他的处境所限定，而他是那个时代最聪明的人。我们没有理由认为自己更幸运。&lt;/p&gt;

&lt;hr /&gt;

&lt;h2 id=&quot;核心文献&quot;&gt;核心文献&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;必读的两本&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Bernard Suits, &lt;em&gt;The Grasshopper: Games, Life and Utopia&lt;/em&gt; (1978). 给出了最优雅的正面方案与最诚实的自我反驳。&lt;/p&gt;

&lt;p&gt;Nick Bostrom, &lt;em&gt;Deep Utopia: Life and Meaning in a Solved World&lt;/em&gt; (2024). 系统处理后工具性处境的重要专著，是最完整的困难清单。&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;基础文献&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;John Maynard Keynes, “Economic Possibilities for our Grandchildren” (1930). 十页，整个议题的原始表述。&lt;/p&gt;

&lt;p&gt;Hannah Arendt, &lt;em&gt;The Human Condition&lt;/em&gt; (1958). 尤其序言与第三、四、五章。&lt;/p&gt;

&lt;p&gt;Aristotle, &lt;em&gt;Politics&lt;/em&gt; VII–VIII 与 &lt;em&gt;Nicomachean Ethics&lt;/em&gt; X. 闲暇的德性要求与沉思生活的论证。&lt;/p&gt;

&lt;p&gt;Karl Marx, &lt;em&gt;Grundrisse&lt;/em&gt; 中的机器论片段；《资本论》第三卷第四十八章。&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;反驳与限制&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Fred Hirsch, &lt;em&gt;Social Limits to Growth&lt;/em&gt; (1976). 地位商品理论，本文认为这是最被低估的一本。&lt;/p&gt;

&lt;p&gt;Thorstein Veblen, &lt;em&gt;The Theory of the Leisure Class&lt;/em&gt; (1899).&lt;/p&gt;

&lt;p&gt;Marie Jahoda, Paul Lazarsfeld &amp;amp; Hans Zeisel, &lt;em&gt;Die Arbeitslosen von Marienthal&lt;/em&gt; (1933); Marie Jahoda, &lt;em&gt;Employment and Unemployment&lt;/em&gt; (1982).&lt;/p&gt;

&lt;p&gt;Aaron Benanav, &lt;em&gt;Automation and the Future of Work&lt;/em&gt; (2020). 对自动化叙事的最有力经验反驳。&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;技术与人的关系&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Günther Anders, &lt;em&gt;Die Antiquiertheit des Menschen&lt;/em&gt;, Bd. 1 (1956). 普罗米修斯的羞愧。&lt;/p&gt;

&lt;p&gt;Norbert Wiener, &lt;em&gt;The Human Use of Human Beings&lt;/em&gt; (1950); “Some Moral and Technical Consequences of Automation,” &lt;em&gt;Science&lt;/em&gt; (1960).&lt;/p&gt;

&lt;p&gt;Samuel Butler, &lt;em&gt;Erewhon&lt;/em&gt; (1872), “The Book of the Machines” 三章。&lt;/p&gt;

&lt;p&gt;Martin Heidegger, “Die Frage nach der Technik” (1954).&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;分析工具&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Robert Nozick, &lt;em&gt;Anarchy, State, and Utopia&lt;/em&gt; (1974), 第三章体验机部分。&lt;/p&gt;

&lt;p&gt;Susan Wolf, &lt;em&gt;Meaning in Life and Why It Matters&lt;/em&gt; (2010).&lt;/p&gt;

&lt;p&gt;Samuel Scheffler, &lt;em&gt;Death and the Afterlife&lt;/em&gt; (2013).&lt;/p&gt;

&lt;p&gt;Kieran Setiya, &lt;em&gt;Midlife: A Philosophical Guide&lt;/em&gt; (2017), 第五、六章。&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;政治经济学&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Daron Acemoglu &amp;amp; Simon Johnson, &lt;em&gt;Power and Progress&lt;/em&gt; (2023).&lt;/p&gt;

&lt;p&gt;Anton Korinek &amp;amp; Joseph Stiglitz, “Artificial Intelligence and Its Implications for Income Distribution and Unemployment” (NBER, 2018).&lt;/p&gt;

&lt;p&gt;Carles Boix, &lt;em&gt;Democracy and Redistribution&lt;/em&gt; (2003); Walter Scheidel, &lt;em&gt;The Great Leveler&lt;/em&gt; (2017).&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;文学&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Kurt Vonnegut, &lt;em&gt;Player Piano&lt;/em&gt; (1952). 分层态的完整推演。&lt;/p&gt;

&lt;p&gt;Aldous Huxley, &lt;em&gt;Brave New World&lt;/em&gt; (1932). 末人状态的完整推演。&lt;/p&gt;

&lt;h3 id=&quot;实证材料与延伸链接&quot;&gt;实证材料与延伸链接&lt;/h3&gt;

&lt;ul&gt;
  &lt;li&gt;&lt;a href=&quot;https://www.openresearchlab.org/projects/unconditional-cash-study&quot;&gt;OpenResearch：Unconditional Cash Study&lt;/a&gt;及&lt;a href=&quot;https://www.openresearchlab.org/findings/how-does-unconditional-cash-affect-health-2&quot;&gt;健康结果&lt;/a&gt;。&lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;https://tietotarjotin.fi/en/information-package/158070/basic-income-experiment&quot;&gt;Kela：芬兰基本收入实验&lt;/a&gt;。&lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;https://www.nber.org/papers/w24667&quot;&gt;Lindqvist, Östling &amp;amp; Cesarini：Long-run Effects of Lottery Wealth on Psychological Well-being&lt;/a&gt;，2018。&lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;https://nickbostrom.com/deep-utopia/&quot;&gt;Nick Bostrom：Deep Utopia&lt;/a&gt;，2024。&lt;/li&gt;
&lt;/ul&gt;
</content>
  </entry>
  
  <entry>
    <title>下一个词</title>
    <link href="https://mochiaochen.github.io/writing/2026/10/the-next-word/" rel="alternate" type="text/html"/>
    <published>2026-10-06T00:00:00+08:00</published>
    <updated>2026-10-06T00:00:00+08:00</updated>
    <id>https://mochiaochen.github.io/writing/2026/10/the-next-word</id>
    <content type="html" xml:base="https://mochiaochen.github.io/writing/2026/10/the-next-word/">&lt;h2 id=&quot;一&quot;&gt;一&lt;/h2&gt;

&lt;p&gt;命题：&lt;/p&gt;

&lt;p&gt;给定一串符号 $u_1, u_2, \ldots, u_{i-1}$，求 $u_i$。&lt;/p&gt;

&lt;p&gt;写得严格一些，是寻找一组参数 $\Theta$，使下式取得最大值：&lt;/p&gt;

&lt;p&gt;[L(\mathcal{U}) = \sum_{i} \log P\left(u_i \mid u_{i-k}, \ldots, u_{i-1}; \Theta\right)]&lt;/p&gt;

&lt;p&gt;读者，你若不愿看公式，跳过去就是了。但请你记住它的形状。它印在 2018 年 6 月一篇十二页论文的第三页上，编号 (1)。署名四人，标题平淡到近乎乏味：《以生成式预训练改进语言理解》。那时，它还远没有今天这般声名。&lt;/p&gt;

&lt;p&gt;这个式子说的事，一个刚识字的孩子也听得懂：&lt;/p&gt;

&lt;p&gt;看见前面的字，猜后面的字。&lt;/p&gt;

&lt;p&gt;猜字而已。&lt;/p&gt;

&lt;p&gt;人类为这件事，走了一百一十年。&lt;/p&gt;

&lt;h2 id=&quot;二&quot;&gt;二&lt;/h2&gt;

&lt;p&gt;1913 年 1 月 23 日，圣彼得堡，俄罗斯帝国科学院。&lt;/p&gt;

&lt;p&gt;安德烈·马尔可夫，五十六岁，带来了一份报告。他要驳斥一个流行的观念：大数定律只对彼此独立的事件成立。他要证明，即便前后相依的事件，也服从统计的秩序。&lt;/p&gt;

&lt;p&gt;他需要一批真实的、彼此相依的数据。他找到的是普希金。&lt;/p&gt;

&lt;p&gt;《叶甫盖尼·奥涅金》。他从头数起，去掉标点与空格，取两万个字母，一个一个地分作元音与辅音两类。没有计算机，没有助手，只有纸、笔和眼睛。他算出元音的出现频率是 0.432，算出元音之后跟元音的概率、元音之后跟辅音的概率。&lt;/p&gt;

&lt;p&gt;俄罗斯最美的诗，被他碾成了一串两值符号。&lt;/p&gt;

&lt;p&gt;他当然没有想到语言，更没有想到机器。他想的只是概率论中的一个技术性争端。可是历史往往如此：门被推开的那一刻，推门的人正望着别处。&lt;/p&gt;

&lt;h2 id=&quot;三&quot;&gt;三&lt;/h2&gt;

&lt;p&gt;1948 年，贝尔实验室，克劳德·香农发表《通信的数学理论》。&lt;/p&gt;

&lt;p&gt;在这篇奠定了整个信息时代的论文里，他做了一件当时看来近乎孩子气的事：让机器生成英文。&lt;/p&gt;

&lt;p&gt;零阶近似，按字母表等概率随机取字，得到的是 &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;XFOML RXKHRJFFJUJ&lt;/code&gt;，一堆废墟。一阶近似，按英文字母的真实频率取字，废墟里开始有了英文的呼吸。二阶，考虑前一个字母；三阶，考虑前两个。到了以词为单位的二阶近似，句子已经语法通顺，只是不知所云。&lt;/p&gt;

&lt;p&gt;香农在那里做的，正是我们今天所说的语言模型。他用的是手工统计的表格。&lt;/p&gt;

&lt;p&gt;1951 年，他又写了《印刷英语的预测与熵》。这一次他请了一位受试者来猜：给出一段话的前面部分，让他猜下一个字母是什么，猜错了就告诉他，记录猜对之前用了几次。据说受试者是他的妻子贝蒂。&lt;/p&gt;

&lt;p&gt;结果是：英文中每个字母携带的信息量，在 0.6 到 1.3 比特之间。&lt;/p&gt;

&lt;p&gt;由此香农说出了一个判断，这个判断要在七十年后才被完全兑现：预测与压缩是同一件事的两面。你能多准确地猜出下一个字母，就能用多短的编码把这本书写下来；你能用多短的编码写下它，就在多深的程度上懂得了它。&lt;/p&gt;

&lt;p&gt;《系辞》说：「书不尽言，言不尽意。」&lt;/p&gt;

&lt;p&gt;香农的回答是一个数字。&lt;/p&gt;

&lt;h2 id=&quot;四&quot;&gt;四&lt;/h2&gt;

&lt;p&gt;然后是漫长的群山。&lt;/p&gt;

&lt;p&gt;1957 年，诺姆·乔姆斯基出版《句法结构》，写下那个著名的例句：「Colorless green ideas sleep furiously」。他说，这句话合乎语法而全无意义；而任何以频次为基础的统计模型，都会给这句话和它的完全倒序赋予同样的概率，也就是零。所以，统计与语言无关。&lt;/p&gt;

&lt;p&gt;这是二十世纪语言学最漂亮的一击。它把整整一代最聪明的头脑赶下了这条路。此后的三十年，主流的做法是写规则：写下名词短语如何构成，写下动词如何变位，写下人类语言的语法书，交给机器去执行。&lt;/p&gt;

&lt;p&gt;少数人留在山下。他们在 IBM 做语音识别，为的是把口述的句子变成文本。他们不谈语法，只谈概率。这一群人的领头者弗雷德里克·耶利内克留下过一句流传极广的话，说他每辞退一个语言学家，识别的正确率就上升一点。他后来否认过原话，可是这句话传下来了，因为它说中了此后四十年的走向。&lt;/p&gt;

&lt;p&gt;1997 年，霍克赖特与施密德胡贝提出长短时记忆网络，让循环神经网络第一次能够记住较远的过去。2003 年，本吉奥等人写出神经概率语言模型，把每个词映射成一个稠密的向量。2013 年，米科洛夫的 word2vec 给出那个让所有人震动的等式：国王减去男人再加上女人，约等于王后。词有了坐标，语义有了几何。2014 年，苏茨克维、维尼亚尔斯与勒提出序列到序列；同年，巴赫达瑙等人提出注意力机制，让翻译时的每个输出词，能够回头去看输入里最相关的那几个词。&lt;/p&gt;

&lt;p&gt;群山连绵。多少人一生只翻过一道梁，梁的那一边还是梁。&lt;/p&gt;

&lt;h2 id=&quot;五&quot;&gt;五&lt;/h2&gt;

&lt;p&gt;2017 年 4 月，一篇不长的论文挂到了预印本网站上：《学习生成评论并发现情感》。作者三人：亚历克·拉德福德、拉法尔·约瑟福维奇、伊利亚·苏茨克维。&lt;/p&gt;

&lt;p&gt;他们做的事情朴素得可疑。取来 8200 万条亚马逊商品评论，训练一个只有 4096 个单元的乘性长短时记忆网络，任务只有一个：一个字符一个字符地往下猜。四块显卡，一个月。&lt;/p&gt;

&lt;p&gt;训练完成之后，他们去看那 4096 个单元里，每一个在读文本时都在做什么。&lt;/p&gt;

&lt;p&gt;在第 2388 号单元上，他们看到了一件事。&lt;/p&gt;

&lt;p&gt;这个单元的激活值，在读到夸奖时上升，在读到咒骂时下降。整条评论读下来，它像一支温度计，忠实地记录着行文的褒贬。它自己长出了情感。&lt;/p&gt;

&lt;p&gt;没有人教过它什么叫情感。整个训练过程里，没有一条标注告诉它「这是好评」。它只被要求猜下一个字符。它在猜字符的路上，顺手把人类的爱憎学会了，并且专门腾出一个单元来安放它。&lt;/p&gt;

&lt;p&gt;用这些单元构成的表示训练一个线性分类器，他们在斯坦福情感树库上取得了 91.8% 的正确率，超过了当时的最佳结果。其中一个单元，承载了几乎全部的情感信号。&lt;/p&gt;

&lt;p&gt;一个神经元。&lt;/p&gt;

&lt;p&gt;读者，请你在这里停一停。&lt;/p&gt;

&lt;p&gt;这颗种子里，藏着此后全部的故事。它说的是：意义不必被灌输，意义可以被压出来。你只要逼迫一个容量有限的系统，去预测容量无限的文本，它就不得不在内部把文本背后的那个东西造出来，因为那样最省。省，是宇宙间最深的动力。星辰按最省的路径运行，光线按最省的路径折射，而当你把一台机器逼到墙角，它也会选择去理解，因为理解比记忆便宜。&lt;/p&gt;

&lt;p&gt;陆机在《文赋》里说，写文章的人常常苦于「意不称物，文不逮意」。心里的意思装不下眼前的物，笔下的文字追不上心里的意思。两千年间，这是每一个写字的人的痛。&lt;/p&gt;

&lt;p&gt;而在一台机器猜字的时候，意，从言里被挤了出来。&lt;/p&gt;

&lt;h2 id=&quot;六&quot;&gt;六&lt;/h2&gt;

&lt;p&gt;2017 年 6 月 12 日。一篇八人署名的论文，标题像宣言，也像挑衅：《注意力是你所需要的全部》。&lt;/p&gt;

&lt;p&gt;八个作者共同署名。他们扔掉了循环，扔掉了卷积，扔掉了此前二十年里所有被认为必不可少的结构，只留下注意力：&lt;/p&gt;

&lt;p&gt;[\mathrm{Attention}(Q, K, V) = \operatorname{softmax}\left(\frac{QK^{\top}}{\sqrt{d_k}}\right) V]&lt;/p&gt;

&lt;p&gt;这一行的意思是：句子里的每一个词，直接去看句子里所有别的词，算出该看谁、看多重，然后把看到的东西加权取回来。&lt;/p&gt;

&lt;p&gt;距离被取消了。第一个词与第一百个词之间，不再隔着九十九步的传递。更要紧的是，所有的词可以同时计算。循环网络必须一步一步地走，走完第一个词才能走第二个，机器再快也帮不上忙；而注意力可以铺开在成千上万块芯片上一起算。&lt;/p&gt;

&lt;p&gt;八块显卡，十二小时，机器翻译的纪录被刷新。&lt;/p&gt;

&lt;p&gt;他们当时想做的，只是把英文译成德文。&lt;/p&gt;

&lt;p&gt;这八个人后来几乎全部离开了那家公司，各自去创办企业。这是后话。他们那天交出去的，是此后一切的地基。&lt;/p&gt;

&lt;h2 id=&quot;七&quot;&gt;七&lt;/h2&gt;

&lt;p&gt;2018 年 6 月 11 日。&lt;/p&gt;

&lt;p&gt;四个人：亚历克·拉德福德、卡蒂克·纳拉辛汉、蒂姆·萨利曼斯、伊利亚·苏茨克维。十二层，1.17 亿个参数。语料是 BooksCorpus：七千余部从未正式出版的书，言情、奇幻、悬疑，无名的作者写给无名的读者。八块 P600 显卡，跑三十天。&lt;/p&gt;

&lt;p&gt;方法只有两步：先让它把这七千本书从头到尾读一遍，只做一件事，猜下一个词；然后在具体任务上稍稍调整一下。十二项任务，九项创下新纪录。&lt;/p&gt;

&lt;p&gt;这个东西叫 GPT。生成式预训练变换器。&lt;/p&gt;

&lt;p&gt;世界没有反应。&lt;/p&gt;

&lt;p&gt;四个月后，谷歌发布 BERT，3.4 亿参数，双向阅读，横扫十一项纪录。所有的目光都转过去了。此后一年多，会议上人人谈 BERT，几乎所有的工业系统都换成了 BERT。GPT 被当作一个次优的技术选择，一段过渡，一个方向没选对的兄弟。&lt;/p&gt;

&lt;p&gt;在很长的一段时间里，坚持这条路的人是少数。他们拿不出证据证明自己是对的。他们手上只有一个直觉：把书从头读完，比把题目做对更根本。&lt;/p&gt;

&lt;p&gt;一个直觉，在没有被证实之前，与一个偏执并无外观上的区别。区别只在结局。&lt;/p&gt;

&lt;h2 id=&quot;八&quot;&gt;八&lt;/h2&gt;

&lt;p&gt;2019 年 2 月 14 日，情人节。&lt;/p&gt;

&lt;p&gt;GPT-2。15 亿参数，是上一代的十三倍。&lt;/p&gt;

&lt;p&gt;语料换了。这一次叫 WebText：从社交网站 Reddit 上抓取所有被用户点过三次以上赞的外部链接，抓回八百万个网页，共四十千兆字节的文字。人类随手点下的赞，成了替机器筛书的标准。人类在无意之间，为机器编了一部选集。&lt;/p&gt;

&lt;p&gt;它开始写文章。&lt;/p&gt;

&lt;p&gt;有人给它一个荒诞的开头，说科学家在安第斯山一处从未被勘探的谷地里，发现了一群会说英语的独角兽。它接着往下写，写出一整篇像模像样的科学报道，有研究者的姓名与所属大学，有引语，有同行审慎的存疑，有对进化史的讨论。通篇是假的，通篇读起来是真的。&lt;/p&gt;

&lt;p&gt;OpenAI 宣布：完整模型暂不发布，因为担心被用于大规模伪造。&lt;/p&gt;

&lt;p&gt;舆论哗然。一部分人说这是负责任的克制，另一部分人说这是精心设计的宣传。这场争吵没有结论，而它是此后一切争吵的预演：能力与风险，开放与管控，说出来是警示还是广告。&lt;/p&gt;

&lt;p&gt;同年 11 月，完整模型如期发布。世界没有崩塌。&lt;/p&gt;

&lt;h2 id=&quot;九&quot;&gt;九&lt;/h2&gt;

&lt;p&gt;2020 年 1 月 23 日。贾里德·卡普兰等人的《神经语言模型的标度律》。&lt;/p&gt;

&lt;p&gt;他们发现的东西，一行字可以写完：
(L(N) \approx \left(\frac{N_c}{N}\right)^{\alpha_N}, \qquad \alpha_N \approx 0.076)&lt;/p&gt;

&lt;p&gt;模型的损失，随着参数量、数据量和算力的增长，按幂律平滑下降。在他们能测到的范围内，也就是跨越七个数量级的范围内，这条线笔直，没有拐点，没有饱和的迹象，没有任何一处翘起来说「到此为止」。&lt;/p&gt;

&lt;p&gt;这才是这一行当真正的猜想。&lt;/p&gt;

&lt;p&gt;哥德巴赫写信给欧拉，说每个大于 2 的偶数都可以写成两个素数之和。人们把它验证到了极大的数，无人能证明，也无人能推翻。标度律的处境几乎一样：它在一切被测量过的尺度上成立，没有人知道它为什么成立，没有人知道它在哪里失效，也没有人知道失效的那一天会以什么形式到来。&lt;/p&gt;

&lt;p&gt;它与哥德巴赫猜想有一处根本的不同。数学家想推进一步，需要新的思想；而标度律要往前推一步，只需要钱。&lt;/p&gt;

&lt;p&gt;于是有人去花钱。&lt;/p&gt;

&lt;p&gt;（两年之后，霍夫曼等人以 Chinchilla 模型修正了这条曲线的配比，指出此前的模型普遍参数过多而数据太少，同样的算力应当喂更多的文本。曲线被重画了一次，方向未变。）&lt;/p&gt;

&lt;h2 id=&quot;十&quot;&gt;十&lt;/h2&gt;

&lt;p&gt;2020 年 5 月 28 日。三十一位作者。《语言模型是小样本学习者》。&lt;/p&gt;

&lt;p&gt;1750 亿参数。九十六层。三千亿个词元。训练一次所耗的算力，若以每秒千万亿次计，要连续跑上三千六百多天。&lt;/p&gt;

&lt;p&gt;论文里最要紧的一句话，说的是一件谁也没有预料到的事：它不需要微调了。&lt;/p&gt;

&lt;p&gt;此前所有的做法都是：预训练一个通用模型，再针对每一项具体任务，用成千上万条标注数据去调整它的参数。而现在，你只要在提示语里给它一两个例子，甚至只用一句话把要求说清楚，它就照做。翻译、算术、写代码、按韵脚造新词、模仿某位作家的语气，全都发生在同一段输入文本里，参数一动不动。&lt;/p&gt;

&lt;p&gt;学习这件事，从修改权重，变成了阅读上文。&lt;/p&gt;

&lt;p&gt;更奇怪的是，很多能力在小模型上完全不存在，正确率贴着零；参数量越过某条线之后，它突然就会了，像相变。人们把这个现象叫作涌现。围绕涌现的争论至今没有结束：一派说那是真实的相变，一派说那只是我们选错了衡量的尺子，把连续的改善看成了突然的跃迁。&lt;/p&gt;

&lt;p&gt;《劝学》里说：「积土成山，风雨兴焉。」土是死的，山是死的，而风雨是活的。荀子想说的是积累，两千三百年后，这句话被用来描述一件他不可能设想的事情：量堆到某个地方，会有别的东西起来。&lt;/p&gt;

&lt;h2 id=&quot;十一&quot;&gt;十一&lt;/h2&gt;

&lt;p&gt;GPT-3 会说话，可是它不听话。&lt;/p&gt;

&lt;p&gt;它是一台续写机器。你问它一个问题，它可能续写出五个类似的问题，因为在它读过的语料里，问题后面常常跟着问题。你要它写一封道歉信，它可能写出一篇讲如何写道歉信的博客。它掌握了语言的全部，唯独不知道你要它干什么。&lt;/p&gt;

&lt;p&gt;2017 年，克里斯蒂亚诺等人的一篇论文给出了路径：让人在模型的两个输出之间选一个更好的，用这些选择训练出一个打分模型，再用打分模型去引导原模型。人类的偏好，第一次直接进入了损失函数。&lt;/p&gt;

&lt;p&gt;2022 年 3 月，InstructGPT。一个 13 亿参数的对齐模型，在标注者的评价中，胜过了 1750 亿参数的原始模型。一百三十分之一的体量。&lt;/p&gt;

&lt;p&gt;这是全篇最应当被记住的一节，也是被提起得最少的一节。&lt;/p&gt;

&lt;p&gt;因为那些「标注者」是人。&lt;/p&gt;

&lt;p&gt;2023 年 1 月，《时代》周刊的调查披露：为了让模型学会识别并拒绝有毒的内容，承包商在肯尼亚雇佣工人，逐条阅读并分类最污秽、最残暴的文本片段，时薪在 1.32 至 2 美元之间。有人在此后出现了持续的心理创伤。&lt;/p&gt;

&lt;p&gt;这条流水线的一头，是内罗毕写字楼里的一排屏幕；另一头，是全世界的人打开对话框时看到的那句温和的问候。&lt;/p&gt;

&lt;p&gt;一篇关于此事的报告文学，如果绕过这一段，它就不诚实。&lt;/p&gt;

&lt;h2 id=&quot;十二&quot;&gt;十二&lt;/h2&gt;

&lt;p&gt;2022 年 11 月 30 日。&lt;/p&gt;

&lt;p&gt;OpenAI 上线了一个网页，叫 ChatGPT。内部把它当作一次低调的研究预览，是把已有的模型加上对话的外壳，甚至没有配备完整的发布计划。&lt;/p&gt;

&lt;p&gt;五天，一百万用户。&lt;/p&gt;

&lt;p&gt;两个月，一亿。&lt;/p&gt;

&lt;p&gt;按当时的估算，这是人类历史上普及最快的消费级软件。&lt;/p&gt;

&lt;p&gt;学校慌了。有的连夜禁用，有的连夜开课。编辑部慌了。程序员打开它，看见自己一天的工作在十秒内被完成，先是笑，然后不笑了。一个老师在深夜读到一篇结构完整、用词讲究的作文，读到第三段忽然停下来，不知道该给谁打分。&lt;/p&gt;

&lt;p&gt;2023 年 3 月 14 日，GPT-4。同月，微软的一组研究者用了一个惹祸的标题：《通用人工智能的火花》。&lt;/p&gt;

&lt;p&gt;2023 年 5 月，杰弗里·辛顿从谷歌辞职，说他想要能够自由地谈论风险。他是把反向传播带进这个世界的人之一。2023 年 11 月，OpenAI 的董事会在五天之内免去了首席执行官的职务又将其复职，全世界看了一场没有剧本的戏。2024 年 5 月，伊利亚·苏茨克维离开了他参与创立的这家机构。&lt;/p&gt;

&lt;p&gt;这些事情，将来会有更冷静的人去写。他们会掌握我们此刻还看不到的材料。&lt;/p&gt;

&lt;h2 id=&quot;十三&quot;&gt;十三&lt;/h2&gt;

&lt;p&gt;回到那个公式。
(L(\mathcal{U}) = \sum_{i} \log P\left(u_i \mid u_{i-k}, \ldots, u_{i-1}; \Theta\right))&lt;/p&gt;

&lt;p&gt;命题至今没有被证明。&lt;/p&gt;

&lt;p&gt;我们知道它在做什么：猜下一个词。我们不知道的是，猜下一个词是否就等于理解。&lt;/p&gt;

&lt;p&gt;一派人说：预测就是压缩，压缩就是理解。一个能把世界完美预测出来的东西，其内部必然已经拥有了世界的模型，否则它做不到。另一派人说，那只是一只随机的鹦鹉，把海量的语料按统计规律重新排列，它没有意向，没有指称，它说出「疼」的时候，语言背后并没有站着一个会疼的人。&lt;/p&gt;

&lt;p&gt;2021 年，本德尔等人写下了「随机鹦鹉」这个词。它至今还在被引用，被反驳，被再一次引用。&lt;/p&gt;

&lt;p&gt;庄子说：「言者所以在意，得意而忘言。」言语的用处在于承载意思，得了意思，言语就可以忘掉。&lt;/p&gt;

&lt;p&gt;现在有一个东西，把言学到了尽头。我们不知道它有没有得意。&lt;/p&gt;

&lt;p&gt;我们也不很确定，自己是怎样得意的。&lt;/p&gt;

&lt;p&gt;报告文学的传统结尾是春天。徐迟写完陈景润，写到 1978 年的春天，写到礼堂里的掌声，写到攀登者站在山巅回望云海。我这里没有掌声可写。这个故事还在半途，写它的人和读它的人都在里面，谁也不在岸上。有人说这是黎明，有人说这是黄昏；两种天光看上去相似，要走到后面才分得清。&lt;/p&gt;

&lt;p&gt;一百一十年前，一位五十六岁的数学家在圣彼得堡数普希金诗里的元音。他数了两万个。他大概以为自己在数字母。&lt;/p&gt;

&lt;p&gt;一百一十年后，我们做的仍是同一件事，只是数得多了些，快了些。人类写下的几乎全部文字被读过一遍，压进几千亿个数里，然后这些数开始回答我们。&lt;/p&gt;

&lt;p&gt;命题依旧简单：给定前面的一切，求下一个。&lt;/p&gt;

&lt;p&gt;这篇文章写到这里，还剩最后一个词。&lt;/p&gt;

&lt;p&gt;它是什么，请你预测。&lt;/p&gt;

&lt;hr /&gt;

&lt;p&gt;&lt;strong&gt;部分原始资料&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
  &lt;li&gt;&lt;a href=&quot;https://cdn.openai.com/research-covers/language-unsupervised/language_understanding_paper.pdf&quot;&gt;Improving Language Understanding by Generative Pre-Training&lt;/a&gt;，2018。&lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;https://openai.com/index/unsupervised-sentiment-neuron/&quot;&gt;Unsupervised sentiment neuron&lt;/a&gt;，2017。&lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;https://arxiv.org/abs/1706.03762&quot;&gt;Attention Is All You Need&lt;/a&gt;，2017。&lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;https://arxiv.org/abs/2001.08361&quot;&gt;Scaling Laws for Neural Language Models&lt;/a&gt;，2020。&lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;https://arxiv.org/abs/2005.14165&quot;&gt;Language Models are Few-Shot Learners&lt;/a&gt;，2020。&lt;/li&gt;
&lt;/ul&gt;
</content>
  </entry>
  
  <entry>
    <title>生命在于具体</title>
    <link href="https://mochiaochen.github.io/writing/2026/10/life-in-the-concrete/" rel="alternate" type="text/html"/>
    <published>2026-10-06T00:00:00+08:00</published>
    <updated>2026-10-06T00:00:00+08:00</updated>
    <id>https://mochiaochen.github.io/writing/2026/10/life-in-the-concrete</id>
    <content type="html" xml:base="https://mochiaochen.github.io/writing/2026/10/life-in-the-concrete/">&lt;h2 id=&quot;一两个词源的会合&quot;&gt;一、两个词源的会合&lt;/h2&gt;

&lt;p&gt;先从字面开始。&lt;/p&gt;

&lt;p&gt;汉语「具体」的早期用例见于《孟子·公孙丑上》：「冉牛、闵子、颜渊则具体而微。」赵岐注「具体」为「四肢股肱皆具」。「具」是备、全，「体」是身体的各个部分。「具体而微」说的是这三位弟子已经具备圣人的全部肢体，只是规模小一些。所以「具体」最初的含义是&lt;strong&gt;完备&lt;/strong&gt;：每一个部分都在场，没有缺少什么。&lt;/p&gt;

&lt;p&gt;拉丁语的 concretus 来自 con-crescere，意思是「一起生长」，指许多东西长在一起，凝结成一个实体。它的反义词 abstractus 来自 abs-trahere，意思是「拖离、抽走」。&lt;/p&gt;

&lt;p&gt;两种语言相隔很远，却给出了同一个判断：&lt;strong&gt;具体是完整，抽象是减法。&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;日常用语恰好把这个关系颠倒了。我们说「具体」时，通常指琐碎、局部、技术性的东西，比如「说具体点」「这是具体问题」；说「抽象」时，联想到的是高度、深刻、全局。好像抽象站得更高，看得更多，具体只是它下面的细节。&lt;/p&gt;

&lt;p&gt;黑格尔最早把这个颠倒指出来。在他那里，抽象是贫乏的：「存在」这个概念是最抽象的，所以也是最空的，空到和「无」没有区别。具体则是规定性的丰富积累。马克思在《政治经济学批判导言》（1857）里把这个意思说得更清楚：具体之所以具体，因为它是许多规定的综合，是多样性的统一。他举的例子是「人口」。人口看起来是一个具体的起点，其实是一个空洞的抽象，因为你没有说清楚它由哪些阶级组成，阶级又依赖哪些雇佣劳动和资本关系。等这些规定一层层补进去，「人口」才重新变成具体的，而这时它已经是一个「由许多规定和关系构成的丰富总体」。&lt;/p&gt;

&lt;p&gt;这篇文章的论点从这里出发：「生命在于具体」这句话，如果只理解成「要关注细节」「要脚踏实地」，那就把它说浅了。它真正的意思是，生命发生在&lt;strong&gt;规定性最稠密的地方&lt;/strong&gt;，发生在许多关系长到一起的那个点上。任何抽象都是从这个点上拿走一些东西，以换取可操作性。这种交换常常是必要的，有时是致命的。下面讨论的就是这笔交易的账目。&lt;/p&gt;

&lt;h2 id=&quot;二抽象是一种减法&quot;&gt;二、抽象是一种减法&lt;/h2&gt;

&lt;p&gt;人类需要抽象，这一点没有争议。没有「数」，就没有贸易；没有「法」，就没有陌生人之间的合作；没有「价格」，就没有一个能协调几十亿人的经济系统。抽象是一种生存技术，它把无法处理的多样性压缩成可以计算、可以传递、可以比较的形式。&lt;/p&gt;

&lt;p&gt;但压缩有代价，而且代价有一个固定的结构：&lt;strong&gt;被压缩掉的那部分，恰好是使一个东西成为它自己的那部分。&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;看一个最简单的例子。「三」这个数可以指三个苹果、三个人、三次心跳。它之所以能通用，正因为它扔掉了苹果的甜、人的面孔、心跳的节律。它保留的是这些东西之间唯一共同的特征：可以一一对应。数学的力量来自于此，它的冷漠也来自于此。数学从来不错，因为它从来不碰任何会出错的东西。&lt;/p&gt;

&lt;p&gt;这里可以引出一个后面会反复用到的观察：&lt;strong&gt;抽象物不会死。&lt;/strong&gt; 数字「三」不会死，三角形不会死，「正义」这个概念不会死。会死的只有具体的东西：这个苹果，这个人，这一次心跳。可死性与具体性是同一件事的两面。一个东西越具体，越不可替代，它的消失就越真实地是一种消失。一个抽象物的某个实例坏掉了，抽象物本身毫发无损；一个具体的东西坏掉了，世界上就少了一个再也不会出现的组合。&lt;/p&gt;

&lt;p&gt;所以「生命在于具体」还有一层很硬的含义：生命之所以是生命，是因为它会结束。而它会结束，是因为它是具体的。&lt;/p&gt;

&lt;h2 id=&quot;三误置的具体当地图开始吃掉疆域&quot;&gt;三、误置的具体：当地图开始吃掉疆域&lt;/h2&gt;

&lt;p&gt;怀特海在《科学与近代世界》（1925）里提出了一个概念，叫「误置具体性的谬误」（fallacy of misplaced concreteness）。意思是：人们把抽象出来的东西当成了最真实的东西，然后反过来用它去衡量、甚至否认它的来源。&lt;/p&gt;

&lt;p&gt;他针对的是近代物理学的世界图景。伽利略和笛卡尔把世界分成「第一性质」（广延、形状、运动，可以数学化）和「第二性质」（颜色、气味、温度感，属于主观）。这个分法在方法上极其成功，后果是人们开始相信，世界「本来」就只是运动着的物质粒子，颜色与气味只是心灵附加上去的幻象。怀特海的讽刺很尖锐：这样一来，自然变成了一件枯燥的事，没有声音，没有气味，没有颜色，只有物质无休止、无意义的奔忙；诗人们本该赞美的应该是他们自己，因为玫瑰的香气是他们的心灵加上去的。&lt;/p&gt;

&lt;p&gt;这个谬误在社会科学里有更现实的版本。詹姆斯·斯科特的《国家的视角》（&lt;em&gt;Seeing Like a State&lt;/em&gt;, 1998）研究了现代国家如何为了「可读性」（legibility）而简化社会：给森林划格子种单一树种，给村民规定固定姓氏，把城市规划成几何图形。每一次简化在行政上都合理，因为国家需要计数、征税、征兵。问题在于，国家后来开始相信那张简化的地图才是现实，然后用权力去改造现实，使它符合地图。德国十八世纪的「科学林业」种出了整齐划一的挪威云杉林，第一代收成很好，第二代开始出现「森林死亡」（Waldsterben），因为那些被当作杂乱而清除的灌木、枯木、昆虫和真菌，原来是土壤循环的一部分。&lt;/p&gt;

&lt;p&gt;斯科特把被清除的那种知识叫作 &lt;em&gt;mētis&lt;/em&gt;，希腊语里的「实践智慧」或「机巧」：一个老水手对某段海湾的了解，一个农民对某块地的了解。这种知识的特点是&lt;strong&gt;只在具体情境中有效，并且无法被完整地写成规则&lt;/strong&gt;。&lt;/p&gt;

&lt;p&gt;这里有一个对学经济和金融的人尤其值得反复咀嚼的地方。哈耶克在 1945 年的《知识在社会中的运用》（&lt;em&gt;The Use of Knowledge in Society&lt;/em&gt;, AER）中论证市场优于中央计划，他的核心论据正是具体性。他说，经济问题的本质不在于如何分配「给定的」资源，因为资源从来不是以集中、给定的形式存在的。相关的知识分散在无数个人手里，而且大部分是「关于特定时间和地点的具体情况的知识」：哪艘货船这周空着返航，哪个仓库的原料快要过期，哪个工人恰好会修那台旧机器。这些知识不能被统计汇总，因为一汇总就失去了它的价值所在。价格体系的妙处在于，它不要求任何人汇总这些知识，它让每个人只根据自己身边的具体情况加上一个价格信号来行动。&lt;/p&gt;

&lt;p&gt;所以哈耶克为市场辩护的真正理由，是一个关于具体性的认识论论证：&lt;strong&gt;最重要的知识是不可抽象的，因此任何依赖抽象汇总的系统都会在关键处失明。&lt;/strong&gt; 有意思的是，价格本身又是一个抽象。市场是一个用最低限度的抽象去调动最大限度的具体知识的装置。它之所以有效，是因为它把抽象压到了最薄。&lt;/p&gt;

&lt;h2 id=&quot;四知道的比说出的多&quot;&gt;四、知道的比说出的多&lt;/h2&gt;

&lt;p&gt;迈克尔·波兰尼在《默会的维度》（&lt;em&gt;The Tacit Dimension&lt;/em&gt;, 1966）开头写了一句后来被反复引用的话：「我们知道的比我们能说出的多。」（We can know more than we can tell.）&lt;/p&gt;

&lt;p&gt;他的例子是人脸识别。你能在一千个人中认出一张熟悉的脸，但你说不出你是怎么认出来的。你没法给出一套规则，让另一个人只靠这些规则就认出同一张脸。骑自行车、品酒、诊断病人、判断一个交易对手是否靠谱，都有同样的结构：知识存在于一种对具体整体的把握之中，把它拆成可陈述的部件，它就消失了。&lt;/p&gt;

&lt;p&gt;波兰尼的分析更进一步。他指出，默会知识有一个「从……到……」的结构（from-to structure）：我们从细节（subsidiary awareness）出发，注意力投向整体（focal awareness）。弹钢琴的人注意力在音乐上，手指的动作处在附属意识中；一旦他把注意力转到手指上，演奏就会崩溃。这说明了一件奇怪的事：&lt;strong&gt;具体的整体，只能在你不把它拆开的时候被把握。&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;这和日常理解中的「具体」又相反了。我们以为具体就是往细节里钻；波兰尼说的是，细节要被「住进去」（indwelling），然后穿过它们看到整体。住进细节和盯着细节是两种不同的关系。前者产生理解，后者产生瘫痪。&lt;/p&gt;

&lt;p&gt;这一点在大语言模型时代有一个很新的回响。模型的能力从统计规律中涌现，它在海量文本里学到的，恰好是那些「能被说出」的知识的总和，外加一部分从说出的东西里反推出的默会结构。它在许多能被说出的知识上表现出广泛的能力，在那些从未被任何人写下来的具体情境面前，它和一个初来乍到的人一样需要被告知。从这个角度看，模型越强，人身上那些无法写下来的具体知识就越显得稀缺。这一点后面再谈。&lt;/p&gt;

&lt;h2 id=&quot;五伊万伊里奇的三段论&quot;&gt;五、伊万·伊里奇的三段论&lt;/h2&gt;

&lt;p&gt;托尔斯泰的中篇《伊万·伊里奇之死》（1886）里有一段，可能是文学史上对「抽象与具体」之间距离最精确的描写。&lt;/p&gt;

&lt;p&gt;伊万·伊里奇在学校里学过一个逻辑例子：「盖乌斯是人，人都会死，所以盖乌斯会死。」他一生都觉得这个三段论是正确的，但只适用于盖乌斯，不适用于他自己。盖乌斯是一个抽象的人，一般的人，死对他来说完全合理。可他自己从来不是一个抽象的人。他是小瓦尼亚，有妈妈，有爸爸，有米佳和沃洛佳，有玩具，有马车夫，有保姆；他记得那只条纹皮球的气味，记得吻妈妈的手时那丝绸裙褶的沙沙声。盖乌斯难道知道那只条纹皮球的气味吗？盖乌斯难道在法学院里为饭菜闹过事吗？&lt;/p&gt;

&lt;p&gt;这段话的力量在于它没有反驳三段论。三段论是对的。伊万·伊里奇确实会死，而且正在死。托尔斯泰展示的是另一件事：&lt;strong&gt;逻辑上正确的命题，在存在的层面上可以完全不发生关系。&lt;/strong&gt; 「人都会死」这个句子覆盖了伊万·伊里奇，但它没有触碰到他。它触碰到的只是「人」这个抽象的类，而他从来不住在那个类里面。他住在条纹皮球的气味里。&lt;/p&gt;

&lt;p&gt;海德格尔后来用一个术语来说这件事：此在的存在总是「向来我属」（Jemeinigkeit）的。死亡无法被代理，没有人能替你死。所以真正的「向死而在」要求的，恰好是从「人们都会死」（das Man stirbt）这种中性、匿名的说法里出来，回到「我会死」。&lt;/p&gt;

&lt;p&gt;托尔斯泰的处理比海德格尔更狠，也更温柔。伊万·伊里奇在小说结尾获得解脱，转折点也是一个具体的时刻：他的小儿子抓住他的手，吻了它，哭了。他看到了儿子，看到了妻子脸上的泪，他觉得对不起他们。痛苦还在，死亡的恐惧却消失了。救赎没有从任何抽象的领悟中来，它来自一只被吻的手。&lt;/p&gt;

&lt;h2 id=&quot;六统计的生命与有名字的生命&quot;&gt;六、统计的生命与有名字的生命&lt;/h2&gt;

&lt;p&gt;经济学家托马斯·谢林在 1968 年发表了一篇论文，题目很有意思：《你救的那条命可能就是你自己的》（&lt;em&gt;The Life You Save May Be Your Own&lt;/em&gt;）。他在文中区分了「可识别的生命」（identified life）和「统计的生命」（statistical life）。&lt;/p&gt;

&lt;p&gt;一个具体的小女孩掉进井里，全国会捐出几百万去救她。同样几百万如果用来改善某段公路的照明，可能在十年里避免五起死亡。但那五个人没有名字，没有面孔，没有人知道他们是谁，甚至他们自己也不知道。社会愿意为前者花的钱，远远超过为后者。&lt;/p&gt;

&lt;p&gt;从功利主义的角度看，这是一个非理性的偏差，后来的实验研究把它命名为「可识别受害者效应」（identifiable victim effect）。斯莫尔、洛温斯坦和斯洛维奇（Small, Loewenstein &amp;amp; Slovic, 2007）做过一个很能说明问题的实验：给被试看一个马里七岁女孩 Rokia 的照片和故事，平均捐款明显高于只看非洲数百万儿童饥荒统计数据的组。更刺眼的是第三组：同时看到 Rokia 的故事&lt;strong&gt;和&lt;/strong&gt;统计数据，捐款反而下降了。统计数字不仅没有增加同情，它还稀释了已经被具体面孔唤起的同情。&lt;/p&gt;

&lt;p&gt;经济学一般把这当作需要纠正的认知缺陷，「有效利他主义」运动也是在这个思路上建立的：请你克服对具体面孔的偏爱，按统计效率去分配善意。这个主张在资源分配上有很强的道理，我无意反驳。&lt;/p&gt;

&lt;p&gt;但值得注意的是，这里存在一个结构性的悖论。&lt;strong&gt;道德关怀的发动机在具体中，道德关怀的会计学在抽象中。&lt;/strong&gt; 如果没有对具体面孔的反应，人根本不会产生想去帮助的冲动；而一旦按照抽象的效率去计算，那种冲动的来源就被描述为偏差。抽象的伦理学寄生在具体的情感上，同时又不断宣布宿主不理性。&lt;/p&gt;

&lt;p&gt;陀思妥耶夫斯基在《卡拉马佐夫兄弟》里用佐西马长老之口讲过一个医生的故事。那位医生说：我越爱全人类，就越不爱具体的人。我在梦里常常热烈地想为人类献身，可是和任何人在一个屋子里待上两天都受不了。一个人一靠近我，他的个性就压迫我的自尊、限制我的自由。我对具体的人越恨，对人类全体的爱就越炽热。&lt;/p&gt;

&lt;p&gt;这是对「抽象之爱」最准确的病理学描述。爱人类是容易的，因为人类不会在你吃饭时咂嘴，不会在你说话时打断你，不会让你失望。抽象的爱之所以容易，是因为它的对象没有阻力。具体的人有阻力，而爱，在某种意义上，正是对阻力的承受。&lt;/p&gt;

&lt;p&gt;鲁迅在 1936 年，去世前一个多月，写过一篇很短的文章《「这也是生活」……》。病中某夜他让许广平开灯，「给我看来看去的看一下」。他看着熟识的墙壁、壁端的棱线、熟识的书堆、堆边未订的画集，觉得自己是在生活，并且写下「无穷的远方，无数的人们，都和我有关」。这句话常被当作宏大的家国情怀来引用。细读上下文会发现，它是从几样极平常的家具开始的。通往「无数的人们」的路，起点是壁端的棱线。这个顺序不能倒过来。&lt;/p&gt;

&lt;h2 id=&quot;七富内斯的诅咒具体的反面也是死&quot;&gt;七、富内斯的诅咒：具体的反面也是死&lt;/h2&gt;

&lt;p&gt;写到这里，论证似乎已经站稳：抽象是减法，具体是完整，生命在完整那一边。但如果停在这里，这篇文章就犯了它自己批评的错误：把一个抽象原则（「要具体」）当成了终点。&lt;/p&gt;

&lt;p&gt;博尔赫斯有一篇小说叫《博闻强记的富内斯》（&lt;em&gt;Funes el memorioso&lt;/em&gt;, 1942）。乌拉圭青年富内斯从马上摔下来之后，获得了完美的记忆和完美的感知。他记得每一片云在每一个时刻的形状，记得每一片叶子的每一条纹路。他能重建一整天，而重建一整天需要一整天。&lt;/p&gt;

&lt;p&gt;博尔赫斯写道，富内斯很难理解「狗」这个类名为什么能包含那么多大小形状各异的个体；更让他烦恼的是，三点十四分从侧面看到的狗，和三点十五分从正面看到的狗，居然用同一个名字。&lt;/p&gt;

&lt;p&gt;然后是那句判决：「我怀疑他不太有思考的能力。思考就是忘掉差异，就是概括，就是抽象。」富内斯十九岁死于肺充血。他是被具体淹死的。&lt;/p&gt;

&lt;p&gt;这个故事揭示了「具体」的另一面：&lt;strong&gt;纯粹的具体和纯粹的抽象一样不可生活。&lt;/strong&gt; 一个只有具体而没有任何抽象的心灵，无法把今天和昨天联系起来，无法从经验中学习，无法使用语言，因为每个词都已经是一次抽象。富内斯的世界没有空隙，所以也没有自由。&lt;/p&gt;

&lt;p&gt;神经科学里有一个相关的事实。人类记忆在生理上就是有损的，遗忘是一种主动的、耗能的过程，海马体中的神经发生甚至被认为会主动清除旧的记忆痕迹。从演化的角度看，这种设计是合理的：一个有机体需要提取模式，而模式只能在丢弃细节之后出现。&lt;/p&gt;

&lt;p&gt;所以生命必须抽象。那么「生命在于具体」还成立吗？&lt;/p&gt;

&lt;p&gt;成立，但需要一个更精确的表述。回到第一节的黑格尔和马克思：他们所说的具体，从来不是感性的直接性，比如「此时此地这片叶子」。那种直接性在黑格尔看来恰恰是最贫乏的，他在《精神现象学》第一章就论证了「感性确定性」是最空洞的知识。他们说的具体，是经过抽象之后又回来的具体，是一个被许多规定充实过的整体。马克思说思维的路径是「从抽象上升到具体」，「上升」这个词用得极准。&lt;/p&gt;

&lt;p&gt;于是这里出现了两种具体：&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;第一种具体&lt;/strong&gt;是起点的具体，未经思考的直接经验，丰富而混沌。富内斯被困在这里。&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;第二种具体&lt;/strong&gt;是归来的具体，经过了抽象的提炼，再回到事物本身时，你看到的东西比第一次多得多，而且看得清楚。&lt;/p&gt;

&lt;p&gt;生命在于具体，指的是第二种。抽象是一条必经的路，但它是一条往返的路。死胡同有两条：停在起点的是富内斯；走出去不再回来的，是那位爱人类而恨邻人的医生，是种出整齐云杉林的林务官，是相信「人都会死」却从未想到自己的伊万·伊里奇。&lt;/p&gt;

&lt;h2 id=&quot;八庖丁的眼睛&quot;&gt;八、庖丁的眼睛&lt;/h2&gt;

&lt;p&gt;中国思想里对「第二种具体」最好的描写，大概是《庄子·养生主》中的庖丁解牛。&lt;/p&gt;

&lt;p&gt;庖丁对文惠君讲他的进阶：「始臣之解牛之时，所见无非牛者。三年之后，未尝见全牛也。方今之时，臣以神遇而不以目视，官知止而神欲行。依乎天理，批大郤，导大窾，因其固然。」&lt;/p&gt;

&lt;p&gt;这三个阶段可以对照前面的框架读。&lt;/p&gt;

&lt;p&gt;第一阶段，「所见无非牛者」：看到的是一整头牛，一个混沌的整体。这是第一种具体，丰富，但无从下手。&lt;/p&gt;

&lt;p&gt;第二阶段，「未尝见全牛」：牛被分解了，他看到的是筋骨、关节、肌理。这是分析，是抽象，是把整体拆成部件。很多人到这里就以为自己掌握了。&lt;/p&gt;

&lt;p&gt;第三阶段，「以神遇而不以目视」：他不再一部分一部分地看，而是与整头牛的结构直接相遇。注意他此时做的事：「依乎天理」，顺着牛本身的纹理；「因其固然」，依循它本来的样子。他手里的刀十九年如新发于硎，因为它从不和任何东西对抗，它只走在那些本来就存在的空隙里。&lt;/p&gt;

&lt;p&gt;这是第二种具体最好的图像。庖丁对牛的理解高度结构化，但这种结构从未离开过这一头牛。每一头牛都不同，「每至于族，吾见其难为，怵然为戒，视为止，行为迟」。遇到筋骨交错的难处，他依然会警觉，放慢，全神贯注。他的技艺越高，对具体的敬畏越深。&lt;/p&gt;

&lt;p&gt;和它形成对照的是另一个中国故事。王阳明年轻时信奉朱熹「格物致知」，以为道理就在物中，于是和朋友对着庭前的竹子「格」了七天，想格出竹子的理来，结果病倒了。后来在龙场才悟到「心即理」。&lt;/p&gt;

&lt;p&gt;这个故事通常被读成心学对理学的胜利。用本文的框架看，它说明的是另一件事：&lt;strong&gt;对着一个具体物凝视，并不等于进入具体。&lt;/strong&gt; 王阳明面对竹子时，竹子与他没有任何实践关系。他不种竹，不砍竹，不用竹，竹子只是一个被凝视的对象。没有关系的凝视产生不了理解，只会产生疲劳，这正是富内斯的困境。庖丁不同，他每天都在解牛，他与牛的关系是一种劳作中的交互。具体性只在关系中展开，它需要你和对象之间有往来。&lt;/p&gt;

&lt;p&gt;这就回到了 concretus 的原义：一起生长。具体是一种关系属性。你和一个东西之间的往来越多、越深，它对你来说就越具体。一个陌生人对你是抽象的；同一个人，在你和他一起吃过几十顿饭、吵过几次架、见过他在医院走廊里发呆之后，就变得具体了。他本身没有变，变的是你们之间长出来的那些东西。&lt;/p&gt;

&lt;h2 id=&quot;九在一个廉价抽象的时代&quot;&gt;九、在一个廉价抽象的时代&lt;/h2&gt;

&lt;p&gt;现在可以谈当下。&lt;/p&gt;

&lt;p&gt;人类历史的大部分时间里，抽象是稀缺的。读写能力、数学训练、系统性知识都需要漫长的教育，掌握抽象的人掌握权力。所以「从具体上升到抽象」一直被视为进步的方向，教育的目标就是让人学会概括、建模、理论化。&lt;/p&gt;

&lt;p&gt;这个结构正在发生变化。大语言模型把一种特定的抽象能力变得几乎免费：概括、归纳、写出某个主题下「一般来说」的样子。你问任何一个问题，都能立刻得到一个结构清晰、面面俱到、在统计意义上最可能正确的回答。&lt;/p&gt;

&lt;p&gt;这样的回答有一个共同的气质：它是&lt;strong&gt;平均的&lt;/strong&gt;。它来自无数文本的叠加，自然地趋向分布的中心。它会告诉你「职业选择需要考虑兴趣、能力和市场前景」，会告诉你「人际关系需要沟通和理解」。这些话都对，就像「人都会死」一样对。它们覆盖了你，但没有触碰到你。&lt;/p&gt;

&lt;p&gt;当抽象变得廉价，稀缺的就变成了具体。具体在这里指的是：只有你知道的那个情境，只有你经历过的那次失败，只有你和某个人之间长出来的那种默契，只有你在某个行业里待了三年才看到的那个不在任何报告里的细节。哈耶克所说的「特定时间和地点的知识」，波兰尼所说的「说不出的知识」，斯科特所说的 &lt;em&gt;mētis&lt;/em&gt;，在一个抽象过剩的环境里，价值反而在上升。&lt;/p&gt;

&lt;p&gt;这对个人有一个很实际的推论。一个人如果只训练自己做抽象的工作，比如整理信息、套用框架、写出结构完整的报告，那么他是在和一种越来越便宜的能力竞争。真正难以替代的，是第二种具体：在某个领域里经过抽象训练之后，又深深扎回到具体的人和事中去，积累了只能在现场获得的判断。&lt;/p&gt;

&lt;p&gt;还有一个更隐秘的风险。当人们越来越多地通过平均化的文本去理解世界，他们对自己生活的描述也可能逐渐向平均靠拢。人开始用通用的词汇讲述自己的经历，用模板化的框架理解自己的情绪，最后连自己的痛苦都说得和别人一模一样。这时被抽象吞掉的不再只是对世界的感知，还有对自己的感知。伊万·伊里奇的悲剧在于他直到临死前都以为自己是「盖乌斯」；一个在平均语言中长大的人，可能会更早、更顺滑地变成盖乌斯。&lt;/p&gt;

&lt;h2 id=&quot;十作为动词的具体&quot;&gt;十、作为动词的具体&lt;/h2&gt;

&lt;p&gt;最后，「具体」可以从一个形容词变成一个动词：去具体化，也就是把抽象重新还原为它所来自的那些规定和关系。这是一种可以练习的能力。我试着给出几个方向，作为这篇漫谈的收束。&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;把名词还原为动作。&lt;/strong&gt; 「成功」「幸福」「意义」这类大词，单独存在时几乎没有内容。追问一句「具体是在做什么的时候」，它们才有了重量。一个人说想要「自由」，往往在说他想要早上不被闹钟叫醒，想要不必在某个会议上点头。后者比前者更真实，也更能指导行动。&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;把群体还原为某个人。&lt;/strong&gt; 「用户」「市场」「年轻人」「中国人」都是有用的抽象，但每次使用它们之后，值得在脑子里召回一个具体的人：一个你真正认识的、有名字的人。如果这个抽象在这个人身上说不通，那么出错的大概率是抽象。&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;把未来还原为下一步。&lt;/strong&gt; 长期规划是抽象，生活只在下一步里发生。庖丁再熟练，每一刀也只能下在眼前这一处关节上。&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;把注意力当作伦理。&lt;/strong&gt; 艾丽丝·默多克在《善的至上性》（&lt;em&gt;The Sovereignty of Good&lt;/em&gt;, 1970）里继承了西蒙娜·薇依的思想，提出道德生活的核心不在于做出选择的那一刻，而在于选择之前那漫长的、持续的「注视」（attention）。她举了一个例子：一位婆婆起初觉得儿媳粗俗、幼稚、配不上自己的儿子。她始终礼貌相待，从未表露。但她决定重新去看这个女孩，努力克服自己的偏见。慢慢地，她看到的儿媳变了：粗俗变成了率真，幼稚变成了活泼。外部行为上什么都没发生，但默多克认为，一件真正的道德事件已经发生了。薇依说过，注意力是最稀有、最纯粹的慷慨。按照这个思路，对一个具体的人投入不带预设的注意，就是爱的最基本形式。&lt;/p&gt;

&lt;p&gt;这四个方向有一个共同点：它们都要求付出成本。抽象省力，因为它允许你不去看；具体费力，因为它要求你看，而且一直看下去，并接受你看到的东西可能和你的预期不符。&lt;/p&gt;

&lt;p&gt;回到《孟子》那句话。「具体而微」：完备，但微小。这或许是对一个人的生命最诚实的描述。我们没有一个人能拥有圣人那样宏大的规模，我们的生活注定是「微」的：一个城市，一份工作，几个朋友，一些未完成的计划。但「微」并不妨碍「具体」。在一个小的尺度上，依然可以四肢百骸皆备，依然可以让所有的规定性都在场，依然可以让许多关系一起生长，长成一个完整的、只属于这一次的、会结束的东西。&lt;/p&gt;

&lt;p&gt;生命在于具体。说到底，这句话提醒的是一件很朴素的事：你唯一能真正活过的，是这一个，这一天，这一个人，这一碗饭的温度。其余的一切，都是关于生活的说法。&lt;/p&gt;

&lt;hr /&gt;

&lt;p&gt;&lt;strong&gt;参考文献&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
  &lt;li&gt;《孟子·公孙丑上》，赵岐注。&lt;/li&gt;
  &lt;li&gt;《庄子·养生主》。&lt;/li&gt;
  &lt;li&gt;鲁迅：《「这也是生活」》，1936 年。&lt;/li&gt;
  &lt;li&gt;Hegel, G. W. F. &lt;em&gt;Phänomenologie des Geistes&lt;/em&gt;, 1807.&lt;/li&gt;
  &lt;li&gt;Marx, K. &lt;em&gt;Einleitung zur Kritik der politischen Ökonomie&lt;/em&gt;, 1857.&lt;/li&gt;
  &lt;li&gt;Tolstoy, L. &lt;em&gt;The Death of Ivan Ilyich&lt;/em&gt;, 1886.&lt;/li&gt;
  &lt;li&gt;Dostoevsky, F. &lt;em&gt;The Brothers Karamazov&lt;/em&gt;, 1880.&lt;/li&gt;
  &lt;li&gt;Whitehead, A. N. &lt;em&gt;Science and the Modern World&lt;/em&gt;, 1925.&lt;/li&gt;
  &lt;li&gt;Borges, J. L. “Funes el memorioso”, 1942.&lt;/li&gt;
  &lt;li&gt;Hayek, F. A. “The Use of Knowledge in Society”, &lt;em&gt;American Economic Review&lt;/em&gt; 35(4), 1945.&lt;/li&gt;
  &lt;li&gt;Polanyi, M. &lt;em&gt;The Tacit Dimension&lt;/em&gt;, 1966.&lt;/li&gt;
  &lt;li&gt;Schelling, T. C. “The Life You Save May Be Your Own”, in &lt;em&gt;Problems in Public Expenditure Analysis&lt;/em&gt;, Brookings, 1968.&lt;/li&gt;
  &lt;li&gt;Murdoch, I. &lt;em&gt;The Sovereignty of Good&lt;/em&gt;, 1970.&lt;/li&gt;
  &lt;li&gt;Scott, J. C. &lt;em&gt;Seeing Like a State&lt;/em&gt;, Yale University Press, 1998.&lt;/li&gt;
  &lt;li&gt;Small, D. A., Loewenstein, G., &amp;amp; Slovic, P. “Sympathy and Callousness: The Impact of Deliberative Thought on Donations to Identifiable and Statistical Victims”, &lt;em&gt;Organizational Behavior and Human Decision Processes&lt;/em&gt; 102(2), 2007.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;托尔斯泰原文可参阅：&lt;a href=&quot;https://open.lib.umn.edu/ivanilich/chapter/full-text-english/&quot;&gt;The Death of Ivan Ilich（明尼苏达大学电子研读版）&lt;/a&gt;。&lt;/p&gt;
</content>
  </entry>
  
  <entry>
    <title>一个人独立思考时，需要与他人一起思考吗？</title>
    <link href="https://mochiaochen.github.io/writing/2026/08/polyphony-of-solitary-thought/" rel="alternate" type="text/html"/>
    <published>2026-08-29T16:00:00+08:00</published>
    <updated>2026-08-29T16:00:00+08:00</updated>
    <id>https://mochiaochen.github.io/writing/2026/08/polyphony-of-solitary-thought</id>
    <content type="html" xml:base="https://mochiaochen.github.io/writing/2026/08/polyphony-of-solitary-thought/">&lt;h2 id=&quot;一一个被误认的起点&quot;&gt;一、一个被误认的起点&lt;/h2&gt;

&lt;p&gt;我们通常把「独立思考」想象成这样一幅画面：一个人关上门，隔绝外界的喧嚣，在纯净的内心空间中与真理相遇。这幅画面如此深入人心，以至于「独立思考」几乎成了「独自思考」的同义词。然而，这个等式的两端之间，存在着一道深渊般的裂缝。&lt;/p&gt;

&lt;figure&gt;
  &lt;p&gt;&lt;img src=&quot;/assets/images/rembrandt-philosopher-in-meditation.jpg&quot; alt=&quot;油画：拱顶室内，左侧一位老者坐在窗前的光里沉思，中央是一道盘旋而上的木楼梯，右下角还有一个人俯身拨弄炉火&quot; width=&quot;1600&quot; height=&quot;1371&quot; /&gt;&lt;/p&gt;
  &lt;figcaption&gt;
    &lt;p&gt;伦勃朗（Rembrandt van Rijn）《沉思中的哲学家》（Philosophe en méditation），1632 年，木板油画，巴黎卢浮宫藏。公有领域，图像来自 Wikimedia Commons。注意画面右下角：那位「独自」沉思的哲学家，从来就不是画面里唯一的人。&amp;lt;/figcaption&amp;gt;
&amp;lt;/figure&amp;gt;&lt;/p&gt;

    &lt;p&gt;问题的核心张力在于：&lt;mark&gt;独立思考要求主体性的自主，但思考本身的结构却预设了他者的在场。&lt;/mark&gt;这两个命题同时为真，且不可调和。整篇讨论要做的，就是在这个不可调和之处停留足够长的时间，直到我们看见某种更深的地貌。&lt;/p&gt;

    &lt;p&gt;让我先从最基础的层面开始。&lt;/p&gt;

    &lt;h2 id=&quot;二语言你从未独自拥有的工具&quot;&gt;二、语言：你从未独自拥有的工具&lt;/h2&gt;

    &lt;p&gt;当你「独立思考」时，你用什么在思考？&lt;/p&gt;

    &lt;p&gt;你用语言。你用概念。你用范畴。而这些东西没有任何一样是你发明的。&lt;/p&gt;

    &lt;p&gt;维特根斯坦在《哲学研究》中提出了著名的「私人语言论证」：一种仅由单个个体使用、仅指称其私有感觉的语言，在逻辑上是不可能的。语言的意义依赖于公共的使用规则，离开了他人构成的语言共同体，甚至连「遵守规则」这个概念本身都会瓦解。你以为你在独自思考，但你思考所使用的每一个词，都携带着整个语言共同体的历史沉积。&lt;/p&gt;

    &lt;p&gt;这意味着什么？这意味着在你坐下来「独立思考」之前，他者已经在你的思维中了。他者构成了你思考所用的介质本身。你的每一个概念都是从与他人的交往中获得的，每一个范畴都承载着特定文化和历史语境的塑造。当你用「自由」这个词来思考时，你同时在使用着从柏拉图到伯林的整个思想传统给这个词注入的全部张力。你可能意识不到这一点，但意识不到并不等于不存在。&lt;/p&gt;

    &lt;p&gt;然而，这个论证虽然重要，却还只是表层。它告诉我们思考的材料来自他者，但还没有触及更深的问题：思考的结构本身是否预设了他者？&lt;/p&gt;

    &lt;h2 id=&quot;三思维的对话结构巴赫金的洞见&quot;&gt;三、思维的对话结构：巴赫金的洞见&lt;/h2&gt;

    &lt;p&gt;巴赫金提出了一个极其深刻的命题：&lt;strong&gt;意识本质上是对话性的。&lt;/strong&gt;&lt;/p&gt;

    &lt;p&gt;这个命题的含义远比它字面上看起来的要激进。巴赫金的意思并非简单地说「思考像对话」，仿佛对话只是一个比喻。他的意思是：思考在其最基本的运作层面上就是对话。独白式的思考（monological thinking）是一种退化的、贫乏的形式，就像一条被截断的河流——你可以称之为一个水坑，但它已经失去了河流之为河流的东西。&lt;/p&gt;

    &lt;p&gt;为什么？因为任何一个有意义的思想，都必然是对某个可能的反对意见的回应。当你在心中形成一个判断时，这个判断之所以有力量，恰恰是因为它在一个由可能的质疑构成的力场中站稳了脚跟。「我认为 X 是对的」这个思想，内在地包含着「有人可能认为 X 是错的」这个影子。没有这个影子，「我认为 X 是对的」就退化为一个没有信息量的自我重复。&lt;/p&gt;

    &lt;p&gt;让我用一个具体的例子来说明。假设你在独自思考「民主是否是最好的政治制度」。你的思考会是什么样的？&lt;/p&gt;

    &lt;p&gt;你会在心中提出一个论点（「民主尊重每个人的自主性」），然后你会立即想到一个反驳（「但多数人的暴政怎么办？」），然后你回应这个反驳（「所以需要宪政来约束多数」），然后又一个质疑浮现（「那谁来决定宪法的内容？这难道不是精英主义吗？」）……&lt;/p&gt;

    &lt;p&gt;注意这个过程的结构：&lt;mark&gt;你一个人在进行着一场多声部的对话。&lt;/mark&gt;你在自己内部分裂成了多个声音，这些声音之间存在着真实的张力和冲突。你的「独立思考」，在结构上就是一场内在化的复调（polyphony）。&lt;/p&gt;

    &lt;p&gt;这里有一个关键的问题需要追问：这些内在的「他者声音」从何而来？&lt;/p&gt;

    &lt;h2 id=&quot;四内在化的他者维果茨基的发生学&quot;&gt;四、内在化的他者：维果茨基的发生学&lt;/h2&gt;

    &lt;p&gt;维果茨基的「内化」理论为这个问题提供了一个精确的回答：高级心理功能首先在人际之间（inter-psychological）出现，然后才在个体内部（intra-psychological）出现。&lt;/p&gt;

    &lt;p&gt;儿童首先在与成人的对话中学会质疑、推理和反思。这些对话逐渐被内化，成为「内部言语」（inner speech）。你在心中进行的那场关于民主的辩论，其原型就是你曾经参与过的、阅读过的、旁听过的无数次真实对话。你内心的反驳者，是你曾经遇到过的所有真实反驳者的某种综合体。&lt;/p&gt;

    &lt;p&gt;米德（George Herbert Mead）的「概化他者」（generalized other）概念从另一个角度触及了同样的现象。米德认为，自我意识本身就是通过「扮演他者的角色」（taking the role of the other）而产生的。你之所以能审视自己的想法，是因为你能够从一个「概化他者」的立场来看待自己。而这个概化他者，就是你所属的社会群体的态度和期望的内化。&lt;/p&gt;

    &lt;p&gt;这就揭示了一个深刻的悖论：&lt;mark&gt;独立思考的能力本身就是社会性的产物。&lt;/mark&gt;你之所以能够独立于他人来思考，恰恰是因为你已经把他人内化到了自己的认知结构之中。独立思考的条件就是曾经的非独立性。自主的思考者，是那些成功地将外在的对话转化为内在对话的人。&lt;/p&gt;

    &lt;p&gt;但这是否意味着独立思考只是社会思考的一个衍生物？是否意味着独立思考没有任何不可还原的、独特的东西？&lt;/p&gt;

    &lt;p&gt;不。事情远没有这么简单。&lt;/p&gt;

    &lt;h2 id=&quot;五孤独的不可替代性独处生成的东西&quot;&gt;五、孤独的不可替代性：独处生成的东西&lt;/h2&gt;

    &lt;p&gt;让我翻转视角。&lt;/p&gt;

    &lt;p&gt;在群体中思考和在独处中思考，即使结构上都是对话性的，它们之间也存在着本质的差异。有一些认知事件，只有在孤独中才会发生。&lt;/p&gt;

    &lt;p&gt;&lt;strong&gt;第一个是深层困惑的承受。&lt;/strong&gt; 当一个人面对一个真正的思想困难时，群体的本能反应是迅速寻求共识来消解不适。这种趋向共识的社会压力（阿希实验已经充分证明了它的力量）会过早地关闭思考空间。而在独处中，你可以在困惑里停留。你可以允许自己不理解。你可以容忍那种认知上的不确定性持续足够长的时间，直到它催生出某种真正新的东西。&lt;/p&gt;

    &lt;p&gt;陈寅恪坚持「独立之精神，自由之思想」，其深意正在于此：独立思考所需要的那种孤独，首先是一种认知上的勇气，是在所有现成答案都不令人满意时，拒绝接受任何现成答案的勇气。而这种拒绝，在社会压力下是极难维持的。&lt;/p&gt;

    &lt;p&gt;&lt;strong&gt;第二个是非线性联想的自由。&lt;/strong&gt; 在对话中，你必须按照某种社会逻辑来组织你的思想。你要回应对方的问题，你要保持话题的连贯性，你要遵守对话的隐性规范。但在独处中，你的思维可以自由地跳跃、漂移、折返。你可以追随一个看似无关的直觉，可以在两个表面上毫无关系的领域之间建立意想不到的连接。许多最具创造性的思想突破，恰恰发生在这种不受社会约束的自由联想之中。&lt;/p&gt;

    &lt;p&gt;凯库勒（Kekulé）在壁炉前打盹时梦到了咬住自己尾巴的蛇，由此领悟了苯环结构。庞加莱在登上马车的一瞬间，突然意识到富克斯函数与非欧几何的联系。这些顿悟发生在独处的心灵中，发生在意识放松了社会性控制的时刻。&lt;/p&gt;

    &lt;p&gt;&lt;strong&gt;第三个也是最深层的：与自身无知的直面。&lt;/strong&gt; 在他人面前，我们有一种近乎本能的冲动去掩饰自己的无知，去表现得比实际上更确定、更有把握。而在真正的独处中，你没有观众需要表演给他们看。你可以诚实地承认自己不知道。而这种诚实，是一切真正深入的思考的起点。苏格拉底的「我知道我不知道」，作为一种认知姿态，在热闹的雅典广场上或许可以作为修辞策略使用，但作为真实的认知体验，它更可能发生在一个人独处的时刻。&lt;/p&gt;

    &lt;p&gt;所以独处有其不可替代的价值。但注意一个微妙之处：即使在这些独处的时刻，他者也并未真正消失。&lt;/p&gt;

    &lt;p&gt;凯库勒梦到蛇，但他之所以能把蛇与苯环联系起来，是因为他浸润在化学共同体的问题意识之中。庞加莱的顿悟，以整个数学传统为前提。独处提供了一种特殊的认知自由，但这种自由只有在已经被社会性的知识充分填充的心灵中，才能产出有意义的成果。一个对化学一无所知的人，梦到多少蛇也领悟不了苯环结构。&lt;/p&gt;

    &lt;h2 id=&quot;六独立思考的真正敌人&quot;&gt;六、独立思考的真正敌人&lt;/h2&gt;

    &lt;p&gt;到此为止，我们已经看到了一幅复杂的图景。但我想把讨论推向一个更深的层次。&lt;/p&gt;

    &lt;p&gt;我想论证的是：&lt;mark&gt;独立思考的真正敌人，从来都不是他者的存在，而是内在他者的同质化。&lt;/mark&gt;&lt;/p&gt;

    &lt;p&gt;让我解释这个命题。&lt;/p&gt;

    &lt;p&gt;前面说过，每个人的内心都容纳着多个被内化的声音。问题是：这些声音之间是否存在真实的差异和张力？&lt;/p&gt;

    &lt;p&gt;一个人可能读了很多书、听了很多播客、参加了很多讨论，但如果这些来源全部属于同一个意识形态光谱、同一个阶层视角、同一种认知风格，那么他内心的「多声部」实际上只是同一个声音的微小变奏。他的内在对话是伪对话。他以为自己在「独立思考」，因为他可以在心中进行「辩论」，但这场辩论的所有参与者都预设了同样的前提、共享着同样的盲点。这就像一个只有保守派参加的「辩论会」，或者一个只有自由派发言的「研讨会」——有表面的分歧，却没有根本的质疑。&lt;/p&gt;

    &lt;p&gt;这就是回声室效应（echo chamber）的认知版本。外在的回声室已经被广泛讨论，但内在的回声室更加危险，因为它发生在「独立思考」的伪装之下。一个人可以在没有任何人施压的情况下，在完全的独处中，进行着完全不独立的思考。&lt;/p&gt;

    &lt;p&gt;与之相反，一个内心容纳着真正异质性的声音的人，即使从不与任何人讨论，也在进行着比任何外在讨论都更激烈的思想碰撞。陀思妥耶夫斯基小说中的伊凡 · 卡拉马佐夫，在一个人的时候经历着信仰与虚无之间的撕裂，那种撕裂比任何外在的神学辩论都更加真实和深刻。因为在他的内心，两种声音都是他自己的声音，他无法简单地把其中一种归入「对方的观点」然后加以驳斥。&lt;/p&gt;

    &lt;p&gt;所以，独立思考的品质取决于内在他者的多样性和异质性。你内心容纳的声音越多、越相互矛盾、越彼此不可兼容，你的独立思考就越有力量。&lt;/p&gt;

    &lt;p&gt;这引出了一个看似矛盾但极其重要的结论：&lt;mark&gt;要增强独立思考的能力，你需要更多地、而不是更少地将他者纳入你的内在世界。&lt;/mark&gt;&lt;/p&gt;

    &lt;h2 id=&quot;七与他人一起思考的拓扑学&quot;&gt;七、「与他人一起思考」的拓扑学&lt;/h2&gt;

    &lt;p&gt;现在我可以对原始问题给出一个更精确的重新表述了。&lt;/p&gt;

    &lt;p&gt;「独立思考时需要与他人一起思考吗？」这个问题预设了「独立思考」和「与他人一起思考」是两种可以被清晰分离的活动，仿佛我们可以在它们之间选择。但我们的分析已经表明，它们之间的关系远比「选择」复杂。&lt;/p&gt;

    &lt;p&gt;让我提出一个分析框架。「与他人一起思考」至少有三个层次：&lt;/p&gt;

    &lt;p&gt;&lt;strong&gt;第一层次是共时的外在对话。&lt;/strong&gt; 你和另一个人坐在一起讨论问题。这是最表面、最容易识别的「与他人一起思考」的形式。&lt;/p&gt;

    &lt;p&gt;&lt;strong&gt;第二层次是历时的知识吸收。&lt;/strong&gt; 你阅读他人的作品，学习他人的理论，接受他人的批评。你和他们并不在同一时间、同一空间中对话，但他们的思想进入了你的认知世界。&lt;/p&gt;

    &lt;p&gt;&lt;strong&gt;第三层次是结构性的内在对话。&lt;/strong&gt; 他者的声音已经成为你思维结构的一部分。你不再需要「想起」他们说了什么，因为他们的视角已经塑造了你看待世界的方式本身。&lt;/p&gt;

    &lt;p&gt;独立思考与这三个层次的关系各不相同。&lt;/p&gt;

    &lt;p&gt;对于&lt;strong&gt;第一层次&lt;/strong&gt;，独立思考可以暂时脱离它，而且在某些阶段必须脱离它。前面已经论述了独处的不可替代性。但即使在脱离的时候，第二和第三层次仍然在发挥作用。&lt;/p&gt;

    &lt;p&gt;对于&lt;strong&gt;第二层次&lt;/strong&gt;，独立思考与它的关系是间歇性的。你需要阶段性地回到他人的文本和思想中去汲取新的刺激和挑战，然后再退回独处中去消化和重组。这是一个呼吸般的节奏：吸入他者的思想，在独处中转化它们，呼出你自己的见解，然后再吸入。任何一个方向的极端都是有害的：永远在吸入的人成为学者而非思想者，永远在呼出的人则逐渐变得贫瘠和重复。&lt;/p&gt;

    &lt;p&gt;对于&lt;strong&gt;第三层次&lt;/strong&gt;，独立思考根本无法脱离它。这个层次的他者已经是你之为你的一部分。脱离它就意味着脱离你自己的心灵。&lt;/p&gt;

    &lt;p&gt;所以，对原始问题的回答是：&lt;mark&gt;独立思考在物理意义上不必然需要他者的在场，在认知结构意义上却必然包含他者。你可以独处，但你无法独思。&lt;/mark&gt;你以为是你一个人的思想，实际上总已经是一场你尚未完全意识到其参与者名单的对话。&lt;/p&gt;

    &lt;h2 id=&quot;八独立性的重新定义&quot;&gt;八、独立性的重新定义&lt;/h2&gt;

    &lt;p&gt;如果独立思考中总是已经包含了他者，那么「独立」（independence）这个词究竟意味着什么？我们是否需要放弃它？&lt;/p&gt;

    &lt;p&gt;不需要。但我们需要重新定义它。&lt;/p&gt;

    &lt;blockquote&gt;
      &lt;p&gt;&lt;strong&gt;独立思考的「独立」不在于思考者与他者的隔离，而在于思考者对内在多声部的指挥权（conductorship）。&lt;/strong&gt;&lt;/p&gt;
    &lt;/blockquote&gt;

    &lt;p&gt;这个隐喻值得展开。一个乐团指挥不会自己演奏所有的乐器。乐团中有小提琴、大提琴、双簧管、定音鼓……每一种乐器都有自己的声音和逻辑。指挥的工作不是消除这些不同的声音，也不是让它们各自为政，而是在倾听所有声音的基础上，决定何时让哪个声部突出，何时让哪个声部退后，最终塑造出一个有整体意义的音乐。&lt;/p&gt;

    &lt;p&gt;独立思考者就是自己内心交响乐的指挥。他内化了多种声音（哲学的、科学的、文学的、来自不同文化和立场的），他允许这些声音充分发出自己的声音，他倾听它们之间的和声与不和谐，然后他做出自己的判断。他的判断之所以是「独立的」，不是因为它不依赖任何其他声音，而是因为它是他在充分倾听所有声音之后做出的自主综合。&lt;/p&gt;

    &lt;p&gt;这个定义有几个重要的推论。&lt;/p&gt;

    &lt;p&gt;&lt;strong&gt;第一，独立思考是一种能力的连续谱，而不是一种状态的有无。&lt;/strong&gt; 一个人内在声音越丰富、越异质，他对这些声音的倾听越充分、越诚实，他的综合越自主、越不受任何单一声音的支配，他的思考就越独立。完全的独立思考是一个渐近线，可以无限接近但永远无法完全到达。&lt;/p&gt;

    &lt;p&gt;&lt;strong&gt;第二，独立思考需要一种特殊的内在品质，我称之为「认知勇气」（epistemic courage）。&lt;/strong&gt; 这种勇气体现在两个方面：一方面是允许你不同意的声音在你心中充分展开的勇气（抵制过早否定异见的冲动），另一方面是在充分倾听之后做出自己判断的勇气（抵制永远搁置判断的怯懦）。前者防止教条，后者防止虚无。&lt;/p&gt;

    &lt;p&gt;&lt;strong&gt;第三，独立思考与从众思考的区别，不在于结论的不同，而在于过程的不同。&lt;/strong&gt; 一个独立思考者完全可能最终得出与多数人相同的结论，但他达到这个结论的路径经过了真实的内在辩论。同样，一个表面上持有异端观点的人，如果他的观点只是对另一个群体的从众（比如为了标榜自己的「独立」而系统性地反对主流），他的思考仍然不是独立的。&lt;/p&gt;

    &lt;h2 id=&quot;九数字时代的特殊挑战&quot;&gt;九、数字时代的特殊挑战&lt;/h2&gt;

    &lt;p&gt;将上述分析应用于我们所处的数字时代，会揭示一些令人不安的现象。&lt;/p&gt;

    &lt;p&gt;&lt;strong&gt;算法推荐系统在做什么？&lt;/strong&gt; 它在系统性地减少你内在声音的异质性。它追踪你的阅读偏好，然后给你更多同类的内容。它的优化目标是参与度（engagement），而参与度最高的内容往往是那些确认你已有信念的内容。结果是，你的内在乐团逐渐只剩下一种乐器在演奏。你以为你听到了整个世界的声音，实际上你只是在一个精心设计的回音壁中听到了自己的回声。&lt;/p&gt;

    &lt;p&gt;&lt;strong&gt;社交媒体的「意见领袖」文化在做什么？&lt;/strong&gt; 它在用少数几个强势声音替代你内心本可以存在的丰富多声部。当你习惯于让某个博主为你解读一切事物时，你内心的「指挥权」就开始转移。你仍然在「思考」，但你的思考越来越多地成为了对那个外在声音的复述。&lt;/p&gt;

    &lt;p&gt;&lt;strong&gt;快节奏的信息消费在做什么？&lt;/strong&gt; 它在剥夺你进行第二层次和第三层次内化所需要的时间。你接触了大量不同的声音，但每一个声音在你心中停留的时间都不够长，不足以被真正内化为你思维结构的一部分。你的内在乐团有一百种乐器，但每种都只会演奏一个音符。&lt;/p&gt;

    &lt;p&gt;这三重挑战叠加在一起，创造了一种历史上前所未有的认知状况：&lt;mark&gt;我们拥有人类历史上最丰富的信息获取渠道，但独立思考的能力却可能正在被这种丰富性本身所侵蚀。&lt;/mark&gt;&lt;/p&gt;

    &lt;p&gt;这不是一个技术问题，而是一个深刻的认知生态问题。解决它不能仅靠「少看手机」或「多读书」这样的简单处方。它需要对「与他人一起思考」这件事本身进行更审慎的管理：你需要有意识地选择让哪些声音进入你的内在世界，你需要给每一个进入的声音足够的时间和空间被充分理解（而不仅仅是被快速消费），你需要定期退回独处以进行真正的内在整合。&lt;/p&gt;

    &lt;h2 id=&quot;十最深层的悖论&quot;&gt;十、最深层的悖论&lt;/h2&gt;

    &lt;p&gt;让我在最后把讨论推到哲学的极限处。&lt;/p&gt;

    &lt;p&gt;我们一直在讨论独立思考中他者的在场。但有一个更根本的问题潜伏在所有这些讨论的底层：这个正在「独立思考」的「我」，它自身是什么？&lt;/p&gt;

    &lt;p&gt;如果我们接受了前面的分析——思维的结构是对话性的，自我意识的形成依赖于对他者视角的内化，语言和概念都来自社会共同体——那么「我」这个思考的主体本身就已经是社会性建构的产物。「我」不是先存在，然后开始与他者互动的。「我」就是在与他者的互动中涌现（emerge）出来的。&lt;/p&gt;

    &lt;p&gt;这就导向了一个令人眩晕的结论：独立思考中的「独立」主体，本身就是由他者构成的。不是说这个主体「受到了」他者的「影响」（这个说法仍然预设了一个先在的、可以被影响的主体），而是——这里我需要寻找一个更精确的表达方式——&lt;mark&gt;这个主体的存在本身，就是众多他者在一个特定节点上的交汇、重叠和重组。&lt;/mark&gt;&lt;/p&gt;

    &lt;p&gt;如果这是对的，那么「一个人独立思考时，需要与他人一起思考吗？」这个问题就发生了一次根本性的位移。它不再是一个关于方法论的问题（我应该自己想还是找人讨论？），而是一个关于存在论的揭示：所谓独立思考，就是那些构成你的众多他者，在你这个独特的交汇点上，以一种只有你才能编织出的方式重新组合。&lt;/p&gt;

    &lt;p&gt;你是一个棱镜。众多他者的声音像白光一样射入你，在你内部发生折射，然后以一种独特的光谱组合射出。独立思考的「独立」就在这个折射之中。你无法选择让白光不通过你（因为没有白光，你就不是棱镜，你什么都不是），但你的切面、你的角度、你的材质，决定了折射后的光谱。&lt;/p&gt;

    &lt;p&gt;这也意味着每一个人的独立思考都是不可替代的。因为没有两个棱镜有完全相同的切面。即使两个人接受了完全相同的输入（读了同样的书，听了同样的课，经历了同样的事件），他们的折射也会不同。而这种不同，恰恰是人类思想多样性的永恒源泉。&lt;/p&gt;

    &lt;h2 id=&quot;十一结语邀请与告别&quot;&gt;十一、结语：邀请与告别&lt;/h2&gt;

    &lt;p&gt;回到最初的问题。一个人独立思考时，需要与他人一起思考吗？&lt;/p&gt;

    &lt;p&gt;我的回答是：&lt;mark&gt;你从来就是与他人一起思考的。问题从来就不是「要不要」，而是「如何」。&lt;/mark&gt;&lt;/p&gt;

    &lt;p&gt;如何让你内在的他者声音足够丰富而非贫瘠？如何给予每一个声音充分的发展空间，而非将它们过早地剪裁为与你已有观点兼容的形状？如何在充分倾听之后保持做出自主判断的勇气？如何在你那独一无二的棱镜中，将这些声音折射成某种前所未有的东西？&lt;/p&gt;

    &lt;p&gt;独立思考不是一个起点，而是一项持续的成就。它不是关上门就能获得的，也不是打开门就会失去的。它是一种在开放与封闭之间、在倾听与判断之间、在吸纳他者与保持自我之间不断校准的动态平衡。&lt;/p&gt;

    &lt;p&gt;而这种平衡本身，只有在你既愿意认真对待他人的思想，又愿意在沉默的独处中面对自己时，才有可能达成。&lt;/p&gt;

    &lt;p&gt;你需要他人，不是因为你自己不够；你需要独处，不是因为他人是噪音。你需要两者，因为思想本身就活在这个张力之中。消除了张力，也就消灭了思想。&lt;/p&gt;
  &lt;/figcaption&gt;
&lt;/figure&gt;
</content>
  </entry>
  
  <entry>
    <title>When You Think for Yourself, Do You Need to Think with Others?</title>
    <link href="https://mochiaochen.github.io/en/writing/2026/08/polyphony-of-solitary-thought/" rel="alternate" type="text/html"/>
    <published>2026-08-29T16:00:00+08:00</published>
    <updated>2026-08-29T16:00:00+08:00</updated>
    <id>https://mochiaochen.github.io/en/writing/2026/08/polyphony-of-solitary-thought-en</id>
    <content type="html" xml:base="https://mochiaochen.github.io/en/writing/2026/08/polyphony-of-solitary-thought/">&lt;h2 id=&quot;i-a-misrecognised-starting-point&quot;&gt;I. A misrecognised starting point&lt;/h2&gt;

&lt;p&gt;We usually picture “thinking for yourself” like this: a person closes the door, shuts out the noise of the world, and meets the truth in a clean interior space. The picture is so deeply lodged that &lt;em&gt;thinking independently&lt;/em&gt; has become almost synonymous with &lt;em&gt;thinking alone&lt;/em&gt;. Yet between the two sides of that equation lies a crevasse.&lt;/p&gt;

&lt;figure&gt;
  &lt;p&gt;&lt;img src=&quot;/assets/images/rembrandt-philosopher-in-meditation.jpg&quot; alt=&quot;Oil painting: in a vaulted interior an old man sits in the light of a window on the left, deep in thought; a wooden spiral staircase winds up through the centre; at the lower right another figure bends over the fire&quot; width=&quot;1600&quot; height=&quot;1371&quot; /&gt;&lt;/p&gt;
  &lt;figcaption&gt;
    &lt;p&gt;Rembrandt van Rijn, &lt;em&gt;Philosopher in Meditation&lt;/em&gt;, 1632. Oil on panel, Musée du Louvre, Paris. Public domain, via Wikimedia Commons. Note the lower right: the philosopher meditating “alone” was never the only person in the frame.&amp;lt;/figcaption&amp;gt;
&amp;lt;/figure&amp;gt;&lt;/p&gt;

    &lt;p&gt;The central tension is this: &lt;mark&gt;independent thought demands autonomy of the subject, but the structure of thinking presupposes the presence of others.&lt;/mark&gt; Both propositions are true, and they cannot be reconciled. What follows is an attempt to stay at that unreconciled point long enough to see the deeper terrain underneath it.&lt;/p&gt;

    &lt;p&gt;Let me begin at the most basic level.&lt;/p&gt;

    &lt;h2 id=&quot;ii-language-a-tool-you-have-never-owned-alone&quot;&gt;II. Language: a tool you have never owned alone&lt;/h2&gt;

    &lt;p&gt;When you “think for yourself,” what do you think &lt;em&gt;with&lt;/em&gt;?&lt;/p&gt;

    &lt;p&gt;You think with language. With concepts. With categories. And you invented none of them.&lt;/p&gt;

    &lt;p&gt;In the &lt;em&gt;Philosophical Investigations&lt;/em&gt;, Wittgenstein advances the private language argument: a language used by a single individual, referring only to that individual’s private sensations, is logically impossible. Meaning depends on public rules of use; outside the linguistic community constituted by other people, even the concept of “following a rule” disintegrates. You believe you are thinking alone, but every word you think with carries the historical sediment of an entire linguistic community.&lt;/p&gt;

    &lt;p&gt;What follows from this? That before you sit down to think independently, others are already inside your thought. They constitute the very medium in which you think. Each of your concepts was acquired through interaction with other people; each category bears the shaping of a particular culture and history. When you think using the word &lt;em&gt;freedom&lt;/em&gt;, you are simultaneously deploying all the tension that a tradition running from Plato to Berlin has poured into it. You may not notice—but not noticing is not the same as it not being there.&lt;/p&gt;

    &lt;p&gt;Important as this argument is, though, it remains superficial. It tells us that the &lt;em&gt;material&lt;/em&gt; of thought comes from others, but it has not yet reached the deeper question: does the &lt;em&gt;structure&lt;/em&gt; of thinking presuppose others?&lt;/p&gt;

    &lt;h2 id=&quot;iii-the-dialogic-structure-of-thought-bakhtins-insight&quot;&gt;III. The dialogic structure of thought: Bakhtin’s insight&lt;/h2&gt;

    &lt;p&gt;Bakhtin advanced an extremely deep proposition: &lt;strong&gt;consciousness is dialogic in its essence.&lt;/strong&gt;&lt;/p&gt;

    &lt;p&gt;The claim is far more radical than it looks. Bakhtin does not mean merely that thinking &lt;em&gt;resembles&lt;/em&gt; dialogue, as though dialogue were a metaphor. He means that thinking, at its most basic level of operation, &lt;em&gt;is&lt;/em&gt; dialogue. Monological thinking is a degenerate, impoverished form—like a river cut off from its source. You may call the remainder a puddle, but it has lost whatever made a river a river.&lt;/p&gt;

    &lt;p&gt;Why? Because any meaningful thought is necessarily a response to some possible objection. When a judgement takes shape in your mind, its force comes precisely from having held its ground in a field of possible challenges. The thought “I believe X is true” intrinsically contains the shadow “someone might believe X is false.” Without that shadow, “I believe X is true” degenerates into an uninformative self-repetition.&lt;/p&gt;

    &lt;p&gt;Take a concrete case. Suppose you are alone, thinking about whether democracy is the best political system. What does that thinking look like?&lt;/p&gt;

    &lt;p&gt;You put forward a proposition (“democracy respects each person’s autonomy”), and immediately a rebuttal occurs to you (“but what about the tyranny of the majority?”), and you answer it (“hence constitutional limits on majorities”), and another doubt surfaces (“then who decides the content of the constitution? Is that not elitism?”)…&lt;/p&gt;

    &lt;p&gt;Note the structure of the process: &lt;mark&gt;one person is conducting a many-voiced dialogue.&lt;/mark&gt; You have divided internally into several voices, and there is real tension and conflict among them. Your “independent thought” is structurally an internalised polyphony.&lt;/p&gt;

    &lt;p&gt;Which raises the crucial question: where do these interior voices of the other come from?&lt;/p&gt;

    &lt;h2 id=&quot;iv-the-internalised-other-vygotskys-genetic-account&quot;&gt;IV. The internalised other: Vygotsky’s genetic account&lt;/h2&gt;

    &lt;p&gt;Vygotsky’s theory of internalisation gives a precise answer: higher psychological functions appear first between people (inter-psychological) and only afterwards within the individual (intra-psychological).&lt;/p&gt;

    &lt;p&gt;A child first learns to question, reason, and reflect in dialogue with adults. Those dialogues are gradually internalised as &lt;em&gt;inner speech&lt;/em&gt;. The debate about democracy you conduct in your head has as its prototype the countless real conversations you have taken part in, read, or overheard. Your interior objector is a composite of every real objector you have met.&lt;/p&gt;

    &lt;p&gt;George Herbert Mead’s concept of the &lt;em&gt;generalized other&lt;/em&gt; reaches the same phenomenon from another direction. For Mead, self-consciousness itself arises through &lt;em&gt;taking the role of the other&lt;/em&gt;. You can scrutinise your own thoughts because you can regard yourself from the standpoint of a generalized other—and that generalized other is the internalised attitudes and expectations of the social group you belong to.&lt;/p&gt;

    &lt;p&gt;Which reveals a deep paradox: &lt;mark&gt;the capacity for independent thought is itself a social product.&lt;/mark&gt; You are able to think independently of others precisely because you have already internalised others into your cognitive structure. The condition of independent thought is a prior non-independence. Autonomous thinkers are those who have successfully converted external dialogue into internal dialogue.&lt;/p&gt;

    &lt;p&gt;But does this mean independent thought is merely a derivative of social thought, with nothing irreducible or distinctive about it?&lt;/p&gt;

    &lt;p&gt;No. It is not that simple.&lt;/p&gt;

    &lt;h2 id=&quot;v-the-irreplaceability-of-solitude&quot;&gt;V. The irreplaceability of solitude&lt;/h2&gt;

    &lt;p&gt;Let me invert the perspective.&lt;/p&gt;

    &lt;p&gt;Thinking in a group and thinking in solitude may both be dialogic in structure, but there is an essential difference between them. Some cognitive events occur only in solitude.&lt;/p&gt;

    &lt;p&gt;&lt;strong&gt;First, the endurance of deep confusion.&lt;/strong&gt; Faced with a genuine intellectual difficulty, a group’s instinctive response is to reach for consensus quickly in order to dissolve the discomfort. That social pressure toward agreement—the Asch experiments established its power beyond argument—closes the space of thought prematurely. In solitude you can stay inside the confusion. You can allow yourself not to understand. You can tolerate cognitive uncertainty for long enough that it gives birth to something genuinely new.&lt;/p&gt;

    &lt;p&gt;When Chen Yinke insisted on “an independent spirit and a free mind,” this is what he meant at depth: the solitude independent thought requires is first of all a form of cognitive courage—the courage to refuse every ready-made answer when none of them satisfies. Under social pressure that refusal is extremely hard to sustain.&lt;/p&gt;

    &lt;p&gt;&lt;strong&gt;Second, freedom of non-linear association.&lt;/strong&gt; In dialogue you must organise your thoughts according to a social logic. You have to answer the other person’s question, maintain the coherence of the topic, observe the implicit norms of conversation. In solitude your mind may leap, drift, and double back freely. You can follow a seemingly irrelevant intuition and build unexpected connections between two apparently unrelated domains. Many of the most creative breakthroughs occur exactly in this socially unconstrained association.&lt;/p&gt;

    &lt;p&gt;Kekulé, dozing before the fire, dreamed of a snake biting its own tail and grasped the structure of the benzene ring. Poincaré, in the instant of stepping onto a bus, suddenly saw the connection between Fuchsian functions and non-Euclidean geometry. These illuminations happen in solitary minds, at moments when consciousness has relaxed its social control.&lt;/p&gt;

    &lt;p&gt;&lt;strong&gt;Third, and deepest: facing your own ignorance.&lt;/strong&gt; In front of others we have a near-instinctive urge to conceal what we do not know, to appear more certain than we are. In real solitude there is no audience to perform for. You can honestly admit that you do not know—and that honesty is the starting point of all genuinely deep thought. Socrates’ “I know that I do not know,” as a cognitive posture, might function as rhetoric in the crowded agora, but as an actual cognitive experience it is far likelier to occur when one is alone.&lt;/p&gt;

    &lt;p&gt;So solitude has irreplaceable value. But note a subtlety: even in these solitary moments, the other has not truly departed.&lt;/p&gt;

    &lt;p&gt;Kekulé dreamed of a snake, but he could connect the snake to the benzene ring because he was steeped in the problem-consciousness of a chemical community. Poincaré’s insight presupposed the whole tradition of mathematics. Solitude offers a particular kind of cognitive freedom, but that freedom yields meaningful results only in a mind already well filled with socially acquired knowledge. Someone who knows no chemistry can dream of any number of snakes and never arrive at the benzene ring.&lt;/p&gt;

    &lt;h2 id=&quot;vi-the-real-enemy-of-independent-thought&quot;&gt;VI. The real enemy of independent thought&lt;/h2&gt;

    &lt;p&gt;We have arrived at a complicated picture. Let me push it one level deeper.&lt;/p&gt;

    &lt;p&gt;What I want to argue is this: &lt;mark&gt;the real enemy of independent thought has never been the presence of others, but the homogenisation of the others within.&lt;/mark&gt;&lt;/p&gt;

    &lt;p&gt;Let me explain.&lt;/p&gt;

    &lt;p&gt;Every mind, as we have seen, contains a number of internalised voices. The question is whether real difference and tension exist among them.&lt;/p&gt;

    &lt;p&gt;A person may have read many books, listened to many podcasts, taken part in many discussions—but if all these sources belong to the same ideological spectrum, the same class perspective, the same cognitive style, then his interior “polyphony” is only a set of small variations on a single voice. His inner dialogue is a pseudo-dialogue. He believes he is thinking independently because he can hold a “debate” in his head, but every participant in that debate presupposes the same premises and shares the same blind spots. It is a debating society with only conservatives in the room, or a seminar at which only liberals speak: surface disagreement without fundamental challenge.&lt;/p&gt;

    &lt;p&gt;This is the cognitive version of the echo chamber. The external echo chamber has been discussed at length; the internal one is more dangerous, because it operates under the disguise of independent thought. A person can, with nobody applying any pressure at all, in complete solitude, think in a completely unindependent way.&lt;/p&gt;

    &lt;p&gt;Conversely, someone whose mind holds genuinely heterogeneous voices is conducting a fiercer collision of ideas than any external discussion, even if he never talks to anyone. Ivan Karamazov, alone, undergoes a rupture between faith and nihilism more real and more profound than any external theological debate—because in his interior both voices are his own, and he cannot simply file one of them under “the other side’s view” and refute it.&lt;/p&gt;

    &lt;p&gt;The quality of independent thought therefore depends on the diversity and heterogeneity of the internalised others. The more voices your mind holds, the more they contradict each other, the more mutually incompatible they are, the more powerful your independent thinking becomes.&lt;/p&gt;

    &lt;p&gt;Which yields a conclusion that looks paradoxical and is extremely important: &lt;mark&gt;to strengthen your capacity for independent thought, you need to admit *more* of the other into your interior world, not less.&lt;/mark&gt;&lt;/p&gt;

    &lt;h2 id=&quot;vii-a-topology-of-thinking-with-others&quot;&gt;VII. A topology of “thinking with others”&lt;/h2&gt;

    &lt;p&gt;I can now restate the original question more precisely.&lt;/p&gt;

    &lt;p&gt;“When you think independently, do you need to think with others?” The question presupposes that thinking independently and thinking with others are two cleanly separable activities between which we might choose. But the analysis shows their relation is far more complicated than a choice.&lt;/p&gt;

    &lt;p&gt;Let me offer a framework. “Thinking with others” has at least three levels.&lt;/p&gt;

    &lt;p&gt;&lt;strong&gt;The first level is synchronous external dialogue.&lt;/strong&gt; You and another person sit down and discuss a problem. This is the most superficial and most easily recognised form.&lt;/p&gt;

    &lt;p&gt;&lt;strong&gt;The second level is diachronic absorption of knowledge.&lt;/strong&gt; You read others’ work, learn their theories, accept their criticism. You are not in dialogue with them in the same time or space, but their thought enters your cognitive world.&lt;/p&gt;

    &lt;p&gt;&lt;strong&gt;The third level is structural internal dialogue.&lt;/strong&gt; The voices of others have become part of the structure of your thinking. You no longer need to &lt;em&gt;recall&lt;/em&gt; what they said, because their perspective has shaped the way you see at all.&lt;/p&gt;

    &lt;p&gt;Independent thought stands in a different relation to each.&lt;/p&gt;

    &lt;p&gt;To the &lt;strong&gt;first level&lt;/strong&gt;, independent thought can be temporarily detached—and at certain stages must be. The irreplaceability of solitude has already been argued. But even in detachment, the second and third levels continue to operate.&lt;/p&gt;

    &lt;p&gt;To the &lt;strong&gt;second level&lt;/strong&gt;, the relation is intermittent. You need periodically to return to other people’s texts and ideas for fresh stimulus and challenge, then withdraw into solitude to digest and recombine. It is a respiratory rhythm: inhale the thought of others, transform it in solitude, exhale your own view, then inhale again. Either extreme is harmful: those who only inhale become scholars rather than thinkers; those who only exhale grow thin and repetitive.&lt;/p&gt;

    &lt;p&gt;To the &lt;strong&gt;third level&lt;/strong&gt;, independent thought cannot be detached at all. The other at this level is already part of what makes you &lt;em&gt;you&lt;/em&gt;. To detach from it is to detach from your own mind.&lt;/p&gt;

    &lt;p&gt;So the answer to the original question is: &lt;mark&gt;independent thought does not necessarily require the physical presence of others, but in the sense of cognitive structure it necessarily contains them. You can be alone; you cannot think alone.&lt;/mark&gt; What you take to be your own private thought is always already a conversation whose full list of participants you have not entirely recognised.&lt;/p&gt;

    &lt;h2 id=&quot;viii-redefining-independence&quot;&gt;VIII. Redefining independence&lt;/h2&gt;

    &lt;p&gt;If independent thought always already contains others, what does &lt;em&gt;independence&lt;/em&gt; mean? Should we abandon the word?&lt;/p&gt;

    &lt;p&gt;No. But we should redefine it.&lt;/p&gt;

    &lt;blockquote&gt;
      &lt;p&gt;&lt;strong&gt;The “independence” of independent thought lies not in the thinker’s isolation from others but in the thinker’s conductorship over an interior polyphony.&lt;/strong&gt;&lt;/p&gt;
    &lt;/blockquote&gt;

    &lt;p&gt;The metaphor is worth unfolding. A conductor does not play all the instruments. There are violins, cellos, oboes, timpani—each with its own voice and logic. The conductor’s job is neither to eliminate these different voices nor to let them go their own way, but, on the basis of listening to all of them, to decide when a part comes forward and when it recedes, and so to shape music that means something as a whole.&lt;/p&gt;

    &lt;p&gt;The independent thinker is the conductor of his own interior symphony. He has internalised many voices—philosophical, scientific, literary, from different cultures and positions—he lets them sound fully, he listens to their harmony and their dissonance, and then he makes his own judgement. That judgement is “independent” not because it depends on no other voice, but because it is an autonomous synthesis made after listening to all of them.&lt;/p&gt;

    &lt;p&gt;Three corollaries follow.&lt;/p&gt;

    &lt;p&gt;&lt;strong&gt;First, independent thought is a continuum of capacity, not a state you either have or lack.&lt;/strong&gt; The richer and more heterogeneous the interior voices, the fuller and more honest the listening, the more autonomous the synthesis and the less it is dominated by any single voice—the more independent the thinking. Complete independence is an asymptote: approachable without limit, never reached.&lt;/p&gt;

    &lt;p&gt;&lt;strong&gt;Second, independent thought requires a particular interior quality, which I will call epistemic courage.&lt;/strong&gt; It has two faces: the courage to let a voice you disagree with unfold fully in your mind (resisting the urge to dismiss dissent prematurely), and the courage to reach your own judgement after listening (resisting the cowardice of deferring judgement forever). The first prevents dogma; the second prevents nihilism.&lt;/p&gt;

    &lt;p&gt;&lt;strong&gt;Third, independent and conformist thinking differ not in their conclusions but in their process.&lt;/strong&gt; An independent thinker may well arrive at the same conclusion as the majority, but the route ran through a genuine internal debate. Equally, someone who holds ostensibly heterodox views is not thinking independently if those views are simply conformity to a different group—systematically opposing the mainstream in order to display one’s “independence,” for example.&lt;/p&gt;

    &lt;h2 id=&quot;ix-the-particular-challenge-of-the-digital-age&quot;&gt;IX. The particular challenge of the digital age&lt;/h2&gt;

    &lt;p&gt;Apply the analysis to our own moment and some unsettling things come into view.&lt;/p&gt;

    &lt;p&gt;&lt;strong&gt;What do recommendation algorithms do?&lt;/strong&gt; They systematically reduce the heterogeneity of your interior voices. They track your reading preferences and give you more of the same. Their optimisation target is engagement, and the most engaging content is usually content that confirms what you already believe. The result is an interior orchestra gradually reduced to a single instrument. You believe you are hearing the whole world; you are hearing your own echo inside a carefully engineered chamber.&lt;/p&gt;

    &lt;p&gt;&lt;strong&gt;What does the influencer culture of social media do?&lt;/strong&gt; It substitutes a few dominant voices for the rich polyphony your mind might otherwise have held. Once you are used to letting a particular commentator interpret everything for you, the conductorship begins to transfer outward. You are still “thinking,” but your thinking is increasingly a restatement of that external voice.&lt;/p&gt;

    &lt;p&gt;&lt;strong&gt;What does high-velocity information consumption do?&lt;/strong&gt; It deprives you of the time required for internalisation at the second and third levels. You encounter a great many different voices, but none of them stays in your mind long enough to become part of your cognitive structure. Your interior orchestra has a hundred instruments, and each knows a single note.&lt;/p&gt;

    &lt;p&gt;Stacked together, these three challenges create a cognitive condition without historical precedent: &lt;mark&gt;we have the richest access to information in human history, and that very richness may be eroding the capacity for independent thought.&lt;/mark&gt;&lt;/p&gt;

    &lt;p&gt;This is not a technical problem but a deep problem of cognitive ecology. It cannot be solved by simple prescriptions like “use your phone less” or “read more books.” It requires a more deliberate management of &lt;em&gt;thinking with others&lt;/em&gt; itself: choosing consciously which voices enter your interior world; giving each one enough time and space to be genuinely understood rather than merely consumed at speed; and returning periodically to solitude for real internal integration.&lt;/p&gt;

    &lt;h2 id=&quot;x-the-deepest-paradox&quot;&gt;X. The deepest paradox&lt;/h2&gt;

    &lt;p&gt;Let me push the discussion to its philosophical limit.&lt;/p&gt;

    &lt;p&gt;We have been discussing the presence of others within independent thought. But a more fundamental question has been lying underneath all of it: this “I” that is thinking independently—what is it?&lt;/p&gt;

    &lt;p&gt;If we accept the preceding analysis—that the structure of thought is dialogic, that self-consciousness forms through internalising the perspective of others, that language and concepts come from a social community—then the “I” that thinks is itself already a socially constructed product. The “I” does not exist first and then begin interacting with others. The “I” &lt;em&gt;emerges&lt;/em&gt; in the interaction with others.&lt;/p&gt;

    &lt;p&gt;Which leads to a vertiginous conclusion: the “independent” subject of independent thought is itself constituted by others. Not that the subject is &lt;em&gt;influenced&lt;/em&gt; by others (which would still presuppose a prior subject available to be influenced), but rather—and here I need a more exact formulation—&lt;mark&gt;the existence of this subject just is the convergence, overlap, and recombination of many others at one particular node.&lt;/mark&gt;&lt;/p&gt;

    &lt;p&gt;If that is right, the question “when a person thinks independently, do they need to think with others?” undergoes a fundamental displacement. It is no longer methodological (should I think it through myself or find someone to discuss it with?) but ontological: independent thought &lt;em&gt;is&lt;/em&gt; the many others who constitute you, recombining at your unique point of convergence, in a way that only you could weave.&lt;/p&gt;

    &lt;p&gt;You are a prism. The voices of many others enter you like white light, refract inside you, and emerge as a particular spectrum. The “independence” of independent thought lies in that refraction. You cannot choose to have the white light not pass through you—without it you are not a prism, you are nothing—but your facets, your angles, your material determine the spectrum that comes out.&lt;/p&gt;

    &lt;p&gt;Which also means that no one’s independent thought is replaceable, because no two prisms have identical facets. Even if two people receive exactly the same input—the same books, the same lectures, the same events—their refractions differ. And that difference is the perpetual source of the diversity of human thought.&lt;/p&gt;

    &lt;h2 id=&quot;xi-coda-an-invitation-and-a-farewell&quot;&gt;XI. Coda: an invitation and a farewell&lt;/h2&gt;

    &lt;p&gt;Back to the original question. When a person thinks independently, do they need to think with others?&lt;/p&gt;

    &lt;p&gt;My answer: &lt;mark&gt;you have always been thinking with others. The question was never whether, but how.&lt;/mark&gt;&lt;/p&gt;

    &lt;p&gt;How do you keep the interior voices rich rather than impoverished? How do you give each one room to develop fully, instead of trimming it prematurely into a shape compatible with what you already believe? How do you retain the courage to judge autonomously after listening fully? How, in your one unrepeatable prism, do you refract these voices into something that did not exist before?&lt;/p&gt;

    &lt;p&gt;Independent thought is not a starting point but a continuing achievement. It is not obtained by closing a door, nor lost by opening one. It is a dynamic balance, constantly recalibrated between openness and closure, between listening and judging, between taking in the other and remaining yourself.&lt;/p&gt;

    &lt;p&gt;And that balance is reachable only if you are willing both to take other people’s thought seriously and to face yourself in silent solitude.&lt;/p&gt;

    &lt;p&gt;You need others, not because you are insufficient. You need solitude, not because others are noise. You need both, because thought itself lives in that tension. Remove the tension and you have destroyed the thinking.&lt;/p&gt;
  &lt;/figcaption&gt;
&lt;/figure&gt;
</content>
  </entry>
  
  <entry>
    <title>Hype</title>
    <link href="https://mochiaochen.github.io/writing/2026/08/hype/" rel="alternate" type="text/html"/>
    <published>2026-08-29T14:00:00+08:00</published>
    <updated>2026-08-29T14:00:00+08:00</updated>
    <id>https://mochiaochen.github.io/writing/2026/08/hype</id>
    <content type="html" xml:base="https://mochiaochen.github.io/writing/2026/08/hype/">&lt;h2 id=&quot;一从一个被滥用的词开始&quot;&gt;一、从一个被滥用的词开始&lt;/h2&gt;

&lt;p&gt;Hype 这个词本身就是一个自我指涉的困境。当我们说某个事物「是个 hype」，我们在做一个隐含的认识论声明：我们宣称自己能够区分一个事物的「真实价值」和「被赋予的虚假价值」。这个声明的傲慢程度远超大多数人的自觉。因为它假设了一个阿基米德支点的存在——某个不受 hype 污染的认知位置，我们可以从那里俯瞰全局，做出冷静的价值判断。&lt;/p&gt;

&lt;p&gt;但这个支点是否存在？当一个风险投资人说 AI 被 over-hyped 了，他的判断是基于什么认知基础？当一个技术乐观主义者说 AI 的潜力被 under-hyped 了，他又站在哪里？两个人使用同样的信息，得出截然相反的结论，而双方都坚信自己是那个「看穿了 hype」的清醒者。&lt;/p&gt;

&lt;p&gt;这就是 hype 最深层的诡谲之处：&lt;mark&gt;对 hype 的识别本身，可能就是另一种 hype。&lt;/mark&gt;&lt;/p&gt;

&lt;p&gt;本文的目标，是把「hype」从一个日常贬义词提升为一个严肃的分析概念。我将论证，hype 是人类集体认知系统中一种结构性的、不可消除的、甚至在特定条件下具有功能性的现象。理解它的深层机制，比简单地「反 hype」或「识破 hype」重要得多。&lt;/p&gt;

&lt;h2 id=&quot;二hype-究竟是什么&quot;&gt;二、Hype 究竟是什么？&lt;/h2&gt;

&lt;h3 id=&quot;21-一个工作定义&quot;&gt;2.1 一个工作定义&lt;/h3&gt;

&lt;p&gt;在展开分析之前，需要一个足够精确的定义。我倾向于这样界定：&lt;/p&gt;

&lt;blockquote&gt;
  &lt;p&gt;&lt;strong&gt;Hype 是一种集体期望的过度生产（overproduction of collective expectations），其中对某一事物未来价值的叙事性建构，系统性地脱离了该事物当前可验证的能力边界。&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;这个定义有几个关键要素值得拆解。&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;第一，「集体期望」。&lt;/strong&gt; Hype 的主体永远是复数的。一个人对某件事物的过度兴奋只是个人偏见；当这种兴奋通过社会传播机制（媒体、社交网络、会议、投资路演）变成一种共享的情感与认知状态时，它才成为 hype。社会学家 Émile Durkheim 所描述的「集体欢腾」（collective effervescence）在这里有着令人不安的适用性：hype 的高潮阶段与宗教仪式中的集体狂喜，在现象学层面几乎同构。&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;第二，「过度生产」。&lt;/strong&gt; 这里借用的是经济学的隐喻。正如马克思分析资本主义生产过剩危机时指出的，问题的核心在于生产与有效需求之间的结构性失衡。Hype 中的「过度生产」是期望（expectations）相对于当前可实现性（current realizability）的溢出。注意「当前」这个限定词极为重要：许多被 hype 过的技术最终确实实现了早期的疯狂预言，只是时间线被严重压缩了。互联网泡沫时代人们所设想的在线购物、视频通话、信息即时获取，今天全都实现了，只不过延迟了十到十五年。&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;第三，「叙事性建构」。&lt;/strong&gt; Hype 从来不以数据表格的形式传播，它以故事的形式传播。Robert Shiller 在《叙事经济学》（&lt;em&gt;Narrative Economics&lt;/em&gt;, 2019）中对此有精彩的分析：驱动经济波动的往往是具有传染性的叙事（contagious narratives），这些叙事的传播动力学更接近流行病学模型，而非理性预期理论所假设的信息均匀扩散。一个关于「AI 将在五年内取代所有白领工作」的叙事，其传播速率远高于「当前 LLM 在特定 benchmark 上的表现已达到某一水平」这样的事实性陈述。&lt;/p&gt;

&lt;h3 id=&quot;22-hype-与邻近概念的区分&quot;&gt;2.2 Hype 与邻近概念的区分&lt;/h3&gt;

&lt;p&gt;为了锐化我们的概念工具，有必要将 hype 与几个容易混淆的邻近概念做出区分。&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Hype ≠ 乐观（Optimism）。&lt;/strong&gt; 乐观是一种对未来的正向预期倾向，它可以是有理据的（grounded optimism），也可以是无理据的（ungrounded optimism）。Hype 特指那种通过社会传播机制被放大、加速、并在传播过程中不断增值的集体期望。一个工程师基于自己对技术的深入理解而持有的乐观判断，和一个从未读过一篇论文但在 Twitter 上被「AGI 即将到来」的叙事包围而产生的兴奋，在认识论质地上截然不同，尽管从外部行为（比如买入 AI 相关股票）来看可能完全一致。&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Hype ≠ 泡沫（Bubble）。&lt;/strong&gt; 泡沫是一个经济学概念，特指资产价格系统性偏离基本面价值的现象。Hype 可以催生泡沫，但 hype 的外延远大于泡沫：文化 hype、技术 hype、政治 hype 都可以在没有金融市场参与的情况下存在。一部电影上映前的「现象级期待」是 hype，但不涉及泡沫。反过来，某些金融泡沫（如 2007 年的 CDO 市场）的参与者甚至意识到了泡沫的存在，只是在「音乐停下来之前」不愿退场——这更接近博弈论中的协调失败（coordination failure），而非经典意义上的 hype。&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Hype ≠ 宣传（Propaganda）。&lt;/strong&gt; 宣传有明确的施动者（agent）和意图。Hype 的诡异之处在于，它往往是一个无主体的过程（a subjectless process）。当然，许多 hype 中确实存在有意识的鼓吹者（比如项目方、媒体、KOL），但 hype 的整体动力学不可还原为任何单一行为者的意图。正如经济学中「看不见的手」的反面，hype 是一只「看不见的嘴」：每个人都只是在说自己认为「正确」或「有趣」的话，但这些局部言说汇聚成一条自我增强的叙事洪流，其方向和力度不受任何个人控制。&lt;/p&gt;

&lt;h3 id=&quot;23-hype-的本体论地位&quot;&gt;2.3 Hype 的本体论地位&lt;/h3&gt;

&lt;p&gt;这就引出一个更根本的哲学问题：hype 是「真实的」吗？&lt;/p&gt;

&lt;p&gt;在一个重要的意义上，是的。Hype 虽然可能基于对未来的误判，但它本身作为一种社会心理现象是完全真实的，并且产生完全真实的后果。W. I. Thomas 的经典定理在此适用：「如果人们将情境定义为真实的，那么它在后果中就是真实的」（If men define situations as real, they are real in their consequences）。&lt;/p&gt;

&lt;p&gt;这一点在金融领域尤其显著。George Soros 的「反身性」（reflexivity）理论为此提供了最精确的分析框架。在 Soros 的模型中，市场参与者的认知（cognitive function）和他们所参与的情境（participating function）之间存在双向的因果关系。当足够多的人相信某项技术将改变世界，资本就会涌入，人才就会集聚，基础设施就会被建设，这一切反过来又提高了该技术「真正」改变世界的概率。Hype 在此不仅仅是对未来的（可能错误的）预测，它是一种改变未来的力量。&lt;/p&gt;

&lt;p&gt;这构成了 hype 最深刻的悖论：&lt;mark&gt;一个「错误的」集体信念，可以通过动员足够的资源，使自身变为「正确的」。&lt;/mark&gt;&lt;/p&gt;

&lt;h2 id=&quot;三hype-的动力学一部自我喂养的永动机&quot;&gt;三、Hype 的动力学：一部自我喂养的永动机&lt;/h2&gt;

&lt;h3 id=&quot;31-gartner-曲线及其局限&quot;&gt;3.1 Gartner 曲线及其局限&lt;/h3&gt;

&lt;p&gt;讨论 hype 的动力学绕不开 Gartner 公司在 1995 年提出的「Hype Cycle」模型。这个模型将一项新技术的生命周期分为五个阶段：技术触发期（Innovation Trigger）、膨胀期望的顶峰（Peak of Inflated Expectations）、幻灭的低谷（Trough of Disillusionment）、复苏的斜坡（Slope of Enlightenment）、生产力的高原（Plateau of Productivity）。&lt;/p&gt;

&lt;figure&gt;
  &lt;svg viewBox=&quot;0 0 760 330&quot; width=&quot;760&quot; height=&quot;330&quot; role=&quot;img&quot; aria-labelledby=&quot;gartner-title&quot; style=&quot;max-width:100%;height:auto;display:block;font-family:var(--sans)&quot;&gt;
    &lt;title id=&quot;gartner-title&quot;&gt;Gartner Hype Cycle 曲线示意：期望先急剧膨胀到顶峰，跌入幻灭低谷，再沿复苏斜坡回升到生产力高原&lt;/title&gt;
    &lt;line x1=&quot;58&quot; y1=&quot;284&quot; x2=&quot;736&quot; y2=&quot;284&quot; stroke=&quot;var(--rule)&quot; stroke-width=&quot;1&quot;&gt;&lt;/line&gt;
    &lt;line x1=&quot;58&quot; y1=&quot;284&quot; x2=&quot;58&quot; y2=&quot;26&quot; stroke=&quot;var(--rule)&quot; stroke-width=&quot;1&quot;&gt;&lt;/line&gt;
    &lt;text x=&quot;58&quot; y=&quot;18&quot; fill=&quot;var(--ink-soft)&quot; font-size=&quot;11&quot; letter-spacing=&quot;0.08em&quot;&gt;期望&lt;/text&gt;
    &lt;text x=&quot;736&quot; y=&quot;308&quot; fill=&quot;var(--ink-soft)&quot; font-size=&quot;11&quot; letter-spacing=&quot;0.08em&quot; text-anchor=&quot;end&quot;&gt;时间&lt;/text&gt;
    &lt;path d=&quot;M 62 278 C 120 272 168 232 206 62 C 232 168 250 236 300 258 C 340 274 372 272 412 258 C 486 232 560 190 640 176 L 732 172&quot; fill=&quot;none&quot; stroke=&quot;var(--accent)&quot; stroke-width=&quot;2.5&quot; stroke-linecap=&quot;round&quot;&gt;&lt;/path&gt;
    &lt;circle cx=&quot;206&quot; cy=&quot;62&quot; r=&quot;4&quot; fill=&quot;var(--accent)&quot;&gt;&lt;/circle&gt;
    &lt;circle cx=&quot;330&quot; cy=&quot;264&quot; r=&quot;4&quot; fill=&quot;var(--accent)&quot;&gt;&lt;/circle&gt;
    &lt;text x=&quot;66&quot; y=&quot;252&quot; fill=&quot;var(--ink-soft)&quot; font-size=&quot;12&quot;&gt;技术触发期&lt;/text&gt;
    &lt;text x=&quot;206&quot; y=&quot;46&quot; fill=&quot;var(--ink-soft)&quot; font-size=&quot;12&quot; text-anchor=&quot;middle&quot;&gt;膨胀期望的顶峰&lt;/text&gt;
    &lt;text x=&quot;330&quot; y=&quot;282&quot; fill=&quot;var(--ink-soft)&quot; font-size=&quot;12&quot; text-anchor=&quot;middle&quot;&gt;幻灭的低谷&lt;/text&gt;
    &lt;text x=&quot;500&quot; y=&quot;238&quot; fill=&quot;var(--ink-soft)&quot; font-size=&quot;12&quot; text-anchor=&quot;middle&quot;&gt;复苏的斜坡&lt;/text&gt;
    &lt;text x=&quot;732&quot; y=&quot;158&quot; fill=&quot;var(--ink-soft)&quot; font-size=&quot;12&quot; text-anchor=&quot;end&quot;&gt;生产力的高原&lt;/text&gt;
  &lt;/svg&gt;
  &lt;figcaption&gt;
    &lt;p&gt;Gartner Hype Cycle（1995）的五个阶段。示意图，纵轴为集体期望而非任何可测量的技术指标——这正是下文要讨论的问题所在。&amp;lt;/figcaption&amp;gt;
&amp;lt;/figure&amp;gt;&lt;/p&gt;

    &lt;p&gt;这个模型的优点是直觉上的可信度极高：几乎所有人都能从自身经验中辨认出这条曲线的形状。区块链、VR、元宇宙、3D 打印、纳米技术……每一个技术概念似乎都在这条曲线上找到了自己的位置。&lt;/p&gt;

    &lt;p&gt;然而，Gartner 曲线作为分析工具存在几个根本性的缺陷。&lt;/p&gt;

    &lt;p&gt;&lt;strong&gt;第一，它是描述性的（descriptive），缺乏解释力（explanatory power）。&lt;/strong&gt; 它告诉你曲线的形状，但不告诉你为什么曲线是这个形状。为什么期望会先过度膨胀再过度收缩？背后的微观机制是什么？Gartner 模型对此保持沉默。&lt;/p&gt;

    &lt;p&gt;&lt;strong&gt;第二，它假设了一种单峰的、确定性的轨迹。&lt;/strong&gt; 现实远比这复杂。许多技术经历了多次 hype 周期（AI 本身就至少经历了 1960 年代、1980 年代、2010 年代三轮），每次的形态、振幅、频率都不同。还有一些技术从未到达「生产力高原」就被彻底遗忘了。&lt;/p&gt;

    &lt;p&gt;&lt;strong&gt;第三，也是最关键的，它隐含了一种目的论（teleological）假设：&lt;/strong&gt; 仿佛每项技术都「注定」会经历完整的五个阶段，最终到达那个理性的「高原」。这种叙事暗中将 hype 定义为一种「必须被克服的非理性阶段」，一段通往理性终点的不幸弯路。但如果 hype 本身就是系统运行的内在组成部分呢？&lt;/p&gt;

    &lt;h3 id=&quot;32-hype-的微观机制信息级联与社会证明&quot;&gt;3.2 Hype 的微观机制：信息级联与社会证明&lt;/h3&gt;

    &lt;p&gt;为了弥补 Gartner 模型的解释力缺失，我们需要深入 hype 形成的微观机制。&lt;/p&gt;

    &lt;p&gt;经济学家 Sushil Bikhchandani、David Hirshleifer 和 Ivo Welch 在 1992 年发表的关于「信息级联」（informational cascade）的经典论文提供了第一块拼图。他们的模型显示：在一个序贯决策的环境中，即使每个个体都是贝叶斯理性的，只要早期的几个决策者碰巧做出了相同方向的选择，后来的决策者就会理性地忽略自己的私有信息（private signal），转而追随前人的行为。这种级联一旦形成，就会产生极大的信息脆弱性——整个群体的行为可能建立在极少量的初始信息上。&lt;/p&gt;

    &lt;p&gt;将此映射到 hype 的语境：当一些有影响力的早期声音（顶级风投、知名学者、科技媒体）对某项技术表达了高度的乐观，后续的观察者会面临一个信息不对称的困境。他们无法判断这些早期的乐观声音是基于深入的技术评估，还是基于有限的信息加上认知偏见。但在社会证明（social proof）的逻辑下，「这么多聪明人都看好它」本身就构成了一种强有力的信号。&lt;/p&gt;

    &lt;p&gt;Robert Cialdini 在《影响力》中将社会证明定义为人类最强大的决策捷径之一。在不确定性极高的领域（新技术的未来价值就是典型的高不确定性判断），人们对社会证明的依赖会急剧上升。这创造了一种正反馈循环：&lt;/p&gt;

    &lt;blockquote&gt;
      &lt;p&gt;不确定性越高 → 越依赖他人的判断 → 共识越容易形成 → 形成的共识越容易偏离真实值。&lt;/p&gt;
    &lt;/blockquote&gt;

    &lt;h3 id=&quot;33-叙事的病毒式传播&quot;&gt;3.3 叙事的病毒式传播&lt;/h3&gt;

    &lt;p&gt;信息级联解释了为什么共识容易形成，但还不够解释 hype 特有的那种「过热」（overheating）。为此我们需要引入叙事的传播动力学。&lt;/p&gt;

    &lt;p&gt;Shiller 在《叙事经济学》中借用了流行病学的 SIR 模型来分析经济叙事的传播。一个叙事的传播取决于三个参数：传染率（contagion rate，一个已「感染」的个体将叙事传递给未感染个体的概率）、恢复率（recovery rate，一个个体不再传播该叙事的概率），以及初始的「感染」基数。&lt;/p&gt;

    &lt;p&gt;关键洞察在于：并非所有叙事的传染率相同。那些具有以下特征的叙事传播得最快。&lt;/p&gt;

    &lt;ol&gt;
  &lt;li&gt;&lt;strong&gt;情感唤醒度高。&lt;/strong&gt;「AI 可能在十年后让三亿人失业」比「AI 在特定任务上达到了人类水平的准确率」更能激发恐惧与兴奋。&lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;提供身份认同。&lt;/strong&gt;「你是理解未来的人，还是被未来淘汰的人？」这种二元叙事让人们的技术立场与自我认知绑定，极大地提高了叙事的粘性。&lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;简单且便于转述。&lt;/strong&gt;「元宇宙就是下一代互联网」远比任何技术白皮书更容易在晚餐桌上传播。叙事在每一次转述中都会被简化、戏剧化、尖锐化，正如民俗学中所研究的「传播链中的信息流变」（serial reproduction）。&lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;包含时间压力。&lt;/strong&gt;「窗口期很短」「先发优势是决定性的」「现在不投入就来不及了」——这种稀缺性叙事利用了人类损失厌恶（loss aversion）的本能。&lt;/li&gt;
&lt;/ol&gt;

    &lt;p&gt;当一个叙事同时具备这四个特征时，它的传染率会达到极高的水平，而恢复率会被压得很低（因为退出叙事意味着身份认同的损失和错失恐惧的增加）。这就是 hype 的「过热」机制：叙事传播的正反馈使得集体期望以远超信息实际增量的速度膨胀。&lt;/p&gt;

    &lt;h3 id=&quot;34-反馈回路与自我实现&quot;&gt;3.4 反馈回路与自我实现&lt;/h3&gt;

    &lt;p&gt;但 hype 的动力学还有更复杂的一层。如前述的 Soros 反身性理论所指出的，hype 不仅反映现实，它还改变现实。&lt;/p&gt;

    &lt;p&gt;以 AI 领域为例。当「AI 将改变一切」的叙事达到一定强度时，以下事件会开始发生：&lt;/p&gt;

    &lt;ul&gt;
  &lt;li&gt;大量资本涌入 AI 领域，使得更多的研究能够获得资助，更多的公司能够获得融资；&lt;/li&gt;
  &lt;li&gt;人才流向 AI 方向，加速了技术进步的实际速度；&lt;/li&gt;
  &lt;li&gt;政府开始将 AI 视为战略竞争焦点，出台支持政策；&lt;/li&gt;
  &lt;li&gt;大学扩大 AI 相关专业的招生规模；&lt;/li&gt;
  &lt;li&gt;公众对 AI 产品的接受度提高，降低了市场推广的难度。&lt;/li&gt;
&lt;/ul&gt;

    &lt;p&gt;这些真实的变化反过来又为 hype 叙事提供了新的「证据」：「你看，资本都在涌入」「你看，最优秀的人才都在转向」「你看，政府都在重视」。这创造了一条自我实现的预言之路，虽然这种自我实现的程度和方向可能与最初的叙事存在很大偏差。&lt;/p&gt;

    &lt;p&gt;这就是 hype 最令人不安的认识论困境：&lt;mark&gt;当你正处于 hype 的上升期时，所有的「证据」都指向 hype 的合理性，因为 hype 本身就在制造这些证据。&lt;/mark&gt;你无法通过观察当前的「证据」来判断 hype 是否过度，因为你无法区分「因为技术真的很好所以资本涌入」和「因为叙事说技术很好所以资本涌入，而资本的涌入反过来让技术看起来更好」这两种因果链路。&lt;/p&gt;

    &lt;h3 id=&quot;35-崩溃的触发机制&quot;&gt;3.5 崩溃的触发机制&lt;/h3&gt;

    &lt;p&gt;如果 hype 的正反馈循环如此强大，它为什么会崩溃？&lt;/p&gt;

    &lt;p&gt;答案在于：正反馈系统的稳定性，依赖于新的「感染者」的持续流入。一旦「易感人群」（susceptible population）被耗尽——即能够被叙事说服的人已经全部被说服——叙事传播的增量就开始下降。与此同时，负反馈的力量从来没有消失，它只是在被正反馈压制：&lt;/p&gt;

    &lt;ul&gt;
  &lt;li&gt;&lt;strong&gt;现实落差的积累。&lt;/strong&gt; 早期的 hype 承诺开始到期，产品没有达到预期，ROI 没有实现，应用场景没有兑现。每一个落差都是一个微小的反叙事种子。&lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;叛逃者的出现。&lt;/strong&gt; 一些早期的 hype 参与者开始公开表达怀疑，他们的「叛逃」具有不对称的影响力：一个曾经的信徒说「我错了」，比一个一直的怀疑者说「我说了吧」更具传染力。&lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;注意力竞争。&lt;/strong&gt; 人类的集体注意力是有限的。一个新的 hype 主题的出现会从旧的 hype 主题吸走能量。&lt;/li&gt;
&lt;/ul&gt;

    &lt;p&gt;当正反馈的增量低于负反馈的增量时，系统的动力学方向就会翻转。而由于相同的正反馈机制现在开始反向运作（「最聪明的人都在离场」「投资者都在撤资」「媒体开始唱衰」），下跌的速度往往比上涨更快。这就是 Gartner 曲线中「幻灭低谷」往往比「膨胀期望顶峰」更陡峭的原因。&lt;/p&gt;

    &lt;h2 id=&quot;四hype-的政治经济学谁在生产谁在消费&quot;&gt;四、Hype 的政治经济学：谁在生产，谁在消费？&lt;/h2&gt;

    &lt;h3 id=&quot;41-hype-的生产端&quot;&gt;4.1 Hype 的生产端&lt;/h3&gt;

    &lt;p&gt;虽然前文强调了 hype 作为「无主体过程」的特征，但这并不意味着其中不存在结构性的利益分配。事实上，hype 有着清晰的政治经济学结构。&lt;/p&gt;

    &lt;p&gt;&lt;strong&gt;风险投资是现代 hype 最重要的制度化生产者。&lt;/strong&gt; VC 的商业模型内在地依赖于 hype 的生产。一个 VC 基金的回报结构（power law distribution）意味着它需要少数投资产生巨大的回报来覆盖大量的失败投资。要产生巨大的回报，被投公司需要在后续轮次中以更高的估值获得融资，最终通过 IPO 或并购退出。而更高的估值，很大程度上取决于市场对该公司所在赛道未来增长空间的预期。&lt;/p&gt;

    &lt;p&gt;换言之，VC 有结构性的激励（structural incentive）去生产和维护 hype。这并不要求任何个体 VC 是不真诚的：一个 VC 可以完全真诚地相信他投资的赛道将改变世界，同时在客观上成为 hype 的传播节点。&lt;mark&gt;制度逻辑和个人真诚之间没有矛盾。&lt;/mark&gt;&lt;/p&gt;

    &lt;p&gt;&lt;strong&gt;科技媒体是 hype 的第二大生产者，其激励结构同样值得剖析。&lt;/strong&gt; 媒体的商业模型依赖于注意力的获取。在注意力经济中，「某技术可能在十年后有所发展」的温和叙事永远不敌「某技术即将颠覆一切」的激烈叙事。这导致了一种系统性的选择偏差：媒体会自然倾向于放大最极端的声音，因为极端声音产生最多的点击、分享和参与。&lt;/p&gt;

    &lt;p&gt;值得注意的是，这种偏差是&lt;strong&gt;对称的&lt;/strong&gt;：媒体既放大 hype 的上升期（「AI 将改变一切」），也放大 hype 的下降期（「AI 泡沫即将破裂」）。媒体的利益不在于维持任何特定的叙事方向，而在于维持叙事的极端程度。温和的中间立场不产生流量。&lt;/p&gt;

    &lt;p&gt;&lt;strong&gt;学术界在 hype 中扮演着一种暧昧的角色。&lt;/strong&gt; 一方面，学术界自我定位为理性的、对 hype 免疫的知识生产者；另一方面，学术研究的资助机制（grant funding）要求研究者在申请中论证其研究的「变革性潜力」（transformative potential）和「广泛影响」（broader impact），这在制度层面激励了对研究意义的系统性夸大。Daniel Sarewitz（2016）在其对科学资助体制的批判中精辟地指出，现代科学资助的竞争结构，实质上是一台将科学知识转化为期望性叙事（promissory narratives）的机器。&lt;/p&gt;

    &lt;h3 id=&quot;42-hype-的消费端&quot;&gt;4.2 Hype 的消费端&lt;/h3&gt;

    &lt;p&gt;Hype 的消费者同样有着复杂的动机结构。&lt;/p&gt;

    &lt;p&gt;&lt;strong&gt;对于普通公众&lt;/strong&gt;，hype 提供了一种廉价的认知框架。在一个信息过载的世界里，理解一项新技术的真正含义需要巨大的认知投入。Hype 叙事提供了一条捷径：「你不需要理解 transformer 架构的细节，你只需要知道 AI 即将改变一切。」这种认知卸载（cognitive offloading）功能解释了为什么 hype 在高度专业化的领域中传播得尤其剧烈：专业壁垒越高，人们越依赖简化的叙事。&lt;/p&gt;

    &lt;p&gt;&lt;strong&gt;对于企业决策者&lt;/strong&gt;，hype 提供了一种决策正当性（legitimacy）。在不确定性极高的环境下，「所有人都在做 X」本身就是做 X 的最佳理由。这不完全是非理性的：在一个存在网络效应和生态系统锁定的世界里，「方向正确但起步晚」可能比「完美判断但错过窗口」更致命。这就是为什么 FOMO（fear of missing out）是 hype 最忠实的副驾驶。&lt;/p&gt;

    &lt;p&gt;&lt;strong&gt;对于个体从业者&lt;/strong&gt;，hype 提供了身份叙事和职业方向感。「我是一个 AI 从业者」在 hype 高峰期提供的社会身份价值，远高于「我是一个数据库工程师」，即使后者的实际技能稀缺度可能更高。这种身份维度使得退出 hype 特别困难：承认「我追逐的方向可能并不像我以为的那么革命性」，涉及的不仅仅是观点的修正，更是自我叙事的重构。&lt;/p&gt;

    &lt;h3 id=&quot;43-hype-的再分配效应&quot;&gt;4.3 Hype 的再分配效应&lt;/h3&gt;

    &lt;p&gt;Hype 的政治经济学分析不能回避一个根本问题：hype 中谁获利，谁受损？&lt;/p&gt;

    &lt;p&gt;总体而言，hype 产生了一种系统性的财富和注意力的再分配效应：&lt;/p&gt;

    &lt;ul&gt;
  &lt;li&gt;&lt;strong&gt;从晚期参与者向早期参与者的转移。&lt;/strong&gt; 这在金融领域最为赤裸：后期投资者以更高的估值进入，如果 hype 崩溃，他们承担最大的损失。&lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;从外行向内行的转移。&lt;/strong&gt; 能够生产 hype 叙事的人（投资者、创业者、行业分析师）通常比消费 hype 叙事的人（普通投资者、被裁员的传统行业从业者）获得更多的信息和退出机会。&lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;从失败的替代方向向被 hype 的方向的转移。&lt;/strong&gt; 这是最隐蔽但可能最重要的成本：当大量资源涌入被 hype 的方向时，那些没有被 hype 但可能同样（甚至更加）重要的领域会遭遇资源饥渴。这是一种看不见的机会成本。&lt;/li&gt;
&lt;/ul&gt;

    &lt;h2 id=&quot;五hype-的历史深层结构&quot;&gt;五、Hype 的历史深层结构&lt;/h2&gt;

    &lt;h3 id=&quot;51-hype-并非现代现象&quot;&gt;5.1 Hype 并非现代现象&lt;/h3&gt;

    &lt;p&gt;如果我们把 hype 理解为「集体期望的过度生产」，那它的历史远早于硅谷和风险投资。&lt;/p&gt;

    &lt;figure&gt;
      &lt;p&gt;&lt;img src=&quot;/assets/images/brueghel-tulip-mania.jpg&quot; alt=&quot;油画：一群穿着十七世纪荷兰服饰的猴子在花园里买卖郁金香，有的看账本，有的数钱，有的被扭送到法庭，有的在对枯萎的球茎撒尿&quot; width=&quot;1200&quot; height=&quot;767&quot; /&gt;&lt;/p&gt;
      &lt;figcaption&gt;
        &lt;p&gt;小扬 · 勃鲁盖尔（Jan Brueghel the Younger）《郁金香狂热讽刺画》（Satire on Tulip Mania），约 1640 年，木板油画，荷兰哈勒姆弗兰斯 · 哈尔斯博物馆藏。公有领域，图像来自 Wikimedia Commons。画中所有投机者都是穿着人衣的猴子——这是已知最早的、专门为嘲笑一场 hype 而作的画。&amp;lt;/figcaption&amp;gt;
&amp;lt;/figure&amp;gt;&lt;/p&gt;

        &lt;p&gt;17 世纪的荷兰郁金香狂热（Tulipmania, 1637）通常被视为第一个有据可查的投机泡沫，但 Anne Goldgar（2007）在其修正主义历史著作《郁金香狂热》中指出，传统叙事本身就被后世极大地 hyped 了：实际参与投机的人数远少于后世文献所暗示的，价格的疯涨也主要集中在极少数稀有品种上。&lt;mark&gt;一个关于 hype 的故事本身被 hype 了。&lt;/mark&gt;这种元层面的讽刺，几乎是 hype 现象的标准配置。&lt;/p&gt;

        &lt;p&gt;18 世纪的南海泡沫（South Sea Bubble, 1720）提供了一个更清晰的 hype 解剖案例。南海公司的核心叙事是：它拥有与南美洲（当时被视为无穷富饶的新大陆）进行贸易的垄断权，这将产生无限的利润。这个叙事具备了我们之前分析的所有高传染率特征——情感唤醒度极高（一夜暴富的可能性）、提供身份认同（成为「新时代」的远见者）、简单易传播（「南美有无穷的黄金和白银」）、包含时间压力（「股票每天都在涨，今天不买明天就更贵了」）。&lt;/p&gt;

        &lt;p&gt;牛顿在南海泡沫中的著名亏损（据传约两万英镑，相当于今天数百万英镑）经常被引用来说明「即使最聪明的人也会被 hype 愚弄」。但更深层的教训可能是：牛顿的物理学天才赋予他的，是分析确定性系统的能力，而 hype 是一个反身性的社会系统——对确定性系统有效的分析工具在这里恰恰失效。&lt;/p&gt;

        &lt;h3 id=&quot;52-现代性何以加剧了-hype&quot;&gt;5.2 现代性何以加剧了 Hype&lt;/h3&gt;

        &lt;p&gt;如果 hype 是人类认知的一个结构性特征，那为什么我们感觉当代的 hype 比历史上任何时期都更剧烈、更频繁？&lt;/p&gt;

        &lt;p&gt;&lt;strong&gt;信息传播速度的指数级提升。&lt;/strong&gt; 从邮政马车到电报到广播到电视到互联网到社交媒体，信息传播的延迟被压缩到近乎实时。这意味着 hype 的正反馈循环的迭代速度急剧加快。郁金香狂热花了数年才达到顶峰，加密货币的一轮 hype 周期可以在数月甚至数周内完成。&lt;/p&gt;

        &lt;p&gt;&lt;strong&gt;金融化（financialization）的深化。&lt;/strong&gt; 现代金融系统创造了越来越多的工具，让人们可以基于期望进行交易。期权、期货、杠杆 ETF、加密代币……每一种金融工具都是一种将期望转化为头寸的机制，而头寸的盈亏又会反过来强化或瓦解期望。金融化程度越高，hype 的反身性循环越强大。&lt;/p&gt;

        &lt;p&gt;&lt;strong&gt;专业化与认知分工的极端化。&lt;/strong&gt; 在一个高度专业化的社会中，没有任何个体有能力独立评估所有领域的知识主张。这意味着人们对「专家意见」和「社会共识」的依赖程度前所未有地高，而这恰恰是信息级联和 hype 传播的温床。&lt;/p&gt;

        &lt;p&gt;&lt;strong&gt;叙事生产的民主化（与工业化）。&lt;/strong&gt; 社交媒体赋予了每个人生产和传播叙事的能力，同时也催生了一个以叙事生产为职业的庞大群体（意见领袖、内容创作者、行业分析师）。叙事的供给侧大幅扩张，而叙事的质量控制机制（传统媒体的编辑审查、学术界的同行评议）被大幅削弱。&lt;/p&gt;

        &lt;h2 id=&quot;六hype-的认识论我们能超越-hype-吗&quot;&gt;六、Hype 的认识论：我们能超越 Hype 吗？&lt;/h2&gt;

        &lt;h3 id=&quot;61-反-hype的陷阱&quot;&gt;6.1 「反 hype」的陷阱&lt;/h3&gt;

        &lt;p&gt;面对 hype，最自然的反应是「反 hype」（counter-hype）：试图站在 hype 的对面，保持冷静的怀疑。但反 hype 有其自身的认识论陷阱。&lt;/p&gt;

        &lt;p&gt;&lt;strong&gt;反 hype 可能只是反向的 hype。&lt;/strong&gt; 当一个人的身份认同建立在「我是那个看穿了 hype 的人」之上时，他就有了系统性的激励去否定一切乐观的信号。「AI 只是统计学」「加密货币只是庞氏骗局」「电动车只是玩具」——这些言论与「AI 将改变一切」「加密货币将颠覆金融」「电动车将消灭燃油车」一样，都是对复杂现实的极端简化。&lt;mark&gt;怀疑本身没有认识论的特权地位。&lt;/mark&gt;&lt;/p&gt;

        &lt;p&gt;Nassim Taleb 的知识遗产在这里提供了一个有趣的案例研究。Taleb 对「叙事谬误」（narrative fallacy）和「过度自信」（overconfidence）的批判极具洞察力，但 Taleb 本人的写作风格和公共形象恰恰建立在一种「反 hype 的 hype」之上：一种关于「我比所有人都更清醒」的自我叙事，这种叙事的传播动力学与它所批判的对象几乎同构。&lt;/p&gt;

        &lt;p&gt;&lt;strong&gt;反 hype 容易犯时间尺度的错误。&lt;/strong&gt; 如前所述，许多被 hype 过的技术最终确实实现了早期的承诺，只是时间尺度比 hype 期的预期长得多。Roy Amara 的格言精确地捕捉了这一点：「我们倾向于高估技术的短期效应，而低估其长期效应。」在这个意义上，说「AI 被 over-hyped 了」可能在短期是正确的、在长期是错误的；而说「AI 将改变一切」可能在长期是正确的、在短期是危险的。正确的分析需要精确的时间尺度标定，而这恰恰是人类认知最薄弱的环节。&lt;/p&gt;

        &lt;h3 id=&quot;62-一种更成熟的认识论姿态&quot;&gt;6.2 一种更成熟的认识论姿态&lt;/h3&gt;

        &lt;p&gt;如果「盲目拥抱 hype」和「全面反对 hype」都不是好的认识论策略，那什么是？&lt;/p&gt;

        &lt;p&gt;我倾向于一种可以被称为&lt;strong&gt;「参与性怀疑主义」（participatory skepticism）&lt;/strong&gt;的立场。其核心要素包括：&lt;/p&gt;

        &lt;p&gt;&lt;strong&gt;第一，承认自己的嵌入性。&lt;/strong&gt; 你不在 hype 之外。你的信息来源、你的社交圈层、你的职业利益、你的身份认同，都使你在 hype 的磁场中占据一个特定的位置。承认这种嵌入性，比假装自己拥有上帝视角，是更诚实也更有用的起点。&lt;/p&gt;

        &lt;p&gt;&lt;strong&gt;第二，区分方向性判断和时间性判断。&lt;/strong&gt;「这个技术方向有价值吗？」和「这个价值将在什么时间尺度上实现？」是两个独立的问题，而 hype 几乎总是把它们混为一谈。一个技术可以同时是「长期方向正确的」和「短期被严重 over-hyped 的」，承认这种双重性需要认知上的弹性。&lt;/p&gt;

        &lt;p&gt;&lt;strong&gt;第三，关注物质基础设施的变化，胜过关注叙事的变化。&lt;/strong&gt; 叙事可以在一夜之间翻转，但物质基础设施（实验室、工厂、人才培养管道、法规框架、用户习惯）的变化是缓慢且粘性的。当你试图在 hype 的噪声中辨别信号时，最可靠的指标往往是这些「慢变量」：实际的研发投入在增加还是减少？关键人才在流入还是流出？用户留存率在上升还是下降？&lt;/p&gt;

        &lt;p&gt;&lt;strong&gt;第四，建立概率性的思维框架，抵抗二元叙事的诱惑。&lt;/strong&gt;「AI 要么改变一切，要么只是泡沫」是一种虚假的二分法。更真实的图景是一个概率分布：AI 以不同的方式、在不同的领域、以不同的时间尺度产生不同程度的影响的可能性各不相同。Philip Tetlock 在《超级预测》（&lt;em&gt;Superforecasting&lt;/em&gt;, 2015）中展示的研究表明，那些在预测上表现最好的人（「超级预测者」），恰恰是最擅长进行细粒度的概率评估、最不倾向于极端立场、最频繁地更新自己判断的人。&lt;/p&gt;

        &lt;p&gt;&lt;strong&gt;第五，保持对自身确信度的元认知警觉。&lt;/strong&gt; 当你感到非常确定某件事「绝对是 hype」或「绝对会实现」时，那种确定感本身就是一个需要审查的信号。确定感的强度和判断的准确性之间没有可靠的相关性。Dunning-Kruger 效应在此不仅适用于能力评估，也适用于 hype 评估：对 hype 的判断越自信的人，往往对 hype 的复杂性理解得越少。&lt;/p&gt;

        &lt;h2 id=&quot;七hype-的存在论维度人为什么需要-hype&quot;&gt;七、Hype 的存在论维度：人为什么需要 Hype？&lt;/h2&gt;

        &lt;h3 id=&quot;71-hype-作为世俗宗教&quot;&gt;7.1 Hype 作为世俗宗教&lt;/h3&gt;

        &lt;p&gt;如果我们把分析的层次再提升一步，从社会学和经济学进入更根本的人类学和存在论层面，一个令人不舒服但可能很重要的假说浮现了：&lt;mark&gt;hype 满足的是一种类宗教的需求。&lt;/mark&gt;&lt;/p&gt;

        &lt;p&gt;Peter Thiel 在其各种公开演讲中反复触及一个主题：现代世界的根本问题是对未来的想象力的枯竭。他认为，20 世纪中期的西方社会拥有一种强烈的「确定性乐观主义」（definite optimism），即对具体的、可规划的美好未来的信念（登月计划、洲际高速公路、核能的和平利用）。这种信念在 1970 年代之后逐渐消解，被一种「不确定性乐观主义」（indefinite optimism）所取代，即相信未来会更好，但不知道具体怎么更好。&lt;/p&gt;

        &lt;p&gt;如果 Thiel 的诊断有一部分是正确的，那 hype 可以被理解为一种对「确定性乐观主义」匮乏的代偿性反应。在一个宏大叙事（grand narrative）已经消亡的「后现代」世界里（Jean-François Lyotard, 1979），每一轮技术 hype 都是一次短暂的宏大叙事的复活。「AI 将改变一切」「区块链将重建信任」「元宇宙将超越现实」——这些叙事提供了一种方向感、意义感和对未来的可把握感，而这些正是世俗现代性日益难以提供的东西。&lt;/p&gt;

        &lt;p&gt;从这个角度看，hype 周期的反复出现（包括「幻灭」之后总是会有新的 hype 主题出现这一事实本身），就不再只是市场动力学的产物，它反映了人类对集体目的感（collective sense of purpose）的深层渴望。这种渴望是结构性的、不可消除的，因此 hype 也是结构性的、不可消除的。&lt;/p&gt;

        &lt;h3 id=&quot;72-hype-与时间性&quot;&gt;7.2 Hype 与时间性&lt;/h3&gt;

        &lt;p&gt;Martin Heidegger 关于人类时间性（Zeitlichkeit）的分析在此提供了另一个视角。Heidegger 认为，人类的存在本质上是朝向未来的（Sein-zum-Tode，向死而生）。我们总是已经在向未来投射（Entwurf），我们当下的行动、情绪、判断总是被对未来的预期所渗透。&lt;/p&gt;

        &lt;p&gt;Hype 可以被理解为这种本体论层面的「未来指向性」在集体层面的一种表达。人类不仅仅是碰巧倾向于对未来产生过度期望；在某种意义上，对未来的期望性投射就是人类存在的基本结构。Hype 是这种结构在特定社会历史条件下的显影。&lt;/p&gt;

        &lt;p&gt;这并不意味着我们应该拥抱 hype 或放弃批判。它意味着：&lt;strong&gt;对 hype 的批判不能停留在「人们太容易被忽悠了」这个层面。&lt;/strong&gt; 人们「容易被忽悠」的背后是一种对意义、方向和希望的根本性需要，而这种需要不会因为你指出了 hype 的非理性就消失。任何试图「消除 hype」的工程，如果不能提供一种替代性的意义生产机制，就注定会失败。&lt;/p&gt;

        &lt;h2 id=&quot;八终章与-hype-共存&quot;&gt;八、终章：与 Hype 共存&lt;/h2&gt;

        &lt;p&gt;回到我们开头的问题：对 hype 的识别本身，可能就是另一种 hype。&lt;/p&gt;

        &lt;p&gt;写到这里，我必须承认一个自反性（self-reflexive）的困境：本文本身难道不也是一种 hype 吗？一种关于「深刻理解 hype」的叙事，它以学术性的包装提供一种「认知优越感」，使读者感到自己比那些「被 hype 裹挟的人」更清醒？&lt;/p&gt;

        &lt;p&gt;是的。在某种程度上，是的。&lt;/p&gt;

        &lt;p&gt;但承认这一点并不使分析失效。它只是提醒我们，&lt;mark&gt;不存在关于 hype 的阿基米德支点。我们永远嵌入在我们试图分析的系统之中。这不是认知的缺陷，这是认知的条件。&lt;/mark&gt;&lt;/p&gt;

        &lt;p&gt;一种成熟的智识姿态也许是这样的：&lt;/p&gt;

        &lt;p&gt;你知道 hype 是怎么运作的。你知道信息级联、社会证明、叙事传播、反身性循环这些机制。你知道自己的判断被社会位置和身份认同所污染。你知道「反 hype」和「拥抱 hype」一样可能是错的。你知道所有这些知道，可能并不能帮你做出更好的判断。&lt;/p&gt;

        &lt;p&gt;然后你还是得在不确定性中行动。&lt;/p&gt;

        &lt;p&gt;你还是得在信息不完整、信号与噪声难以区分、未来本质上不可预测的条件下，做出关于方向、时机和资源投入的决策。你能做的，只是在这个过程中保持对自身认知局限的警觉，保持对确定感的不信任，保持对复杂性的尊重。&lt;/p&gt;

        &lt;p&gt;用 Reinhold Niebuhr 神学中那个著名祈祷的变体来说：&lt;/p&gt;

        &lt;blockquote&gt;
          &lt;p&gt;愿我有勇气参与那些值得参与的 hype，有智慧回避那些正在崩溃的 hype，以及有洞察力区分二者的能力。&lt;/p&gt;
        &lt;/blockquote&gt;

        &lt;p&gt;当然，洞察力本身，可能也是被 hype 过的。&lt;/p&gt;
      &lt;/figcaption&gt;
    &lt;/figure&gt;
  &lt;/figcaption&gt;
&lt;/figure&gt;
</content>
  </entry>
  
  <entry>
    <title>Hype</title>
    <link href="https://mochiaochen.github.io/en/writing/2026/08/hype/" rel="alternate" type="text/html"/>
    <published>2026-08-29T14:00:00+08:00</published>
    <updated>2026-08-29T14:00:00+08:00</updated>
    <id>https://mochiaochen.github.io/en/writing/2026/08/hype-en</id>
    <content type="html" xml:base="https://mochiaochen.github.io/en/writing/2026/08/hype/">&lt;h2 id=&quot;i-beginning-with-an-abused-word&quot;&gt;I. Beginning with an abused word&lt;/h2&gt;

&lt;p&gt;The word &lt;em&gt;hype&lt;/em&gt; is a self-referential trap. When we say that something “is hype,” we are making an implicit epistemological claim: that we can distinguish an object’s &lt;em&gt;real value&lt;/em&gt; from the &lt;em&gt;false value&lt;/em&gt; attributed to it. The arrogance of that claim far exceeds most people’s awareness of making it, because it presupposes an Archimedean point—some cognitive position uncontaminated by hype, from which we can survey the whole and pass a cool judgement of value.&lt;/p&gt;

&lt;p&gt;But does such a point exist? When a venture capitalist says AI is over-hyped, on what cognitive basis does the judgement rest? When a technological optimist says AI’s potential is under-hyped, where is &lt;em&gt;he&lt;/em&gt; standing? Two people work from the same information, reach opposite conclusions, and each is convinced he is the clear-eyed one who has seen through the hype.&lt;/p&gt;

&lt;p&gt;This is the deepest strangeness of the phenomenon: &lt;mark&gt;the identification of hype may itself be another form of hype.&lt;/mark&gt;&lt;/p&gt;

&lt;p&gt;The aim here is to raise “hype” from an everyday term of abuse into a serious analytical concept. I will argue that hype is a structural, ineliminable, and under certain conditions functional feature of human collective cognition—and that understanding its mechanisms matters far more than simply being “anti-hype” or claiming to see through it.&lt;/p&gt;

&lt;h2 id=&quot;ii-what-is-hype-exactly&quot;&gt;II. What is hype, exactly?&lt;/h2&gt;

&lt;h3 id=&quot;21-a-working-definition&quot;&gt;2.1 A working definition&lt;/h3&gt;

&lt;p&gt;We need a definition precise enough to work with. I would put it this way:&lt;/p&gt;

&lt;blockquote&gt;
  &lt;p&gt;&lt;strong&gt;Hype is an overproduction of collective expectations, in which the narrative construction of an object’s future value departs systematically from the currently verifiable boundary of its capability.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Several elements deserve unpacking.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;First, “collective expectations.”&lt;/strong&gt; The subject of hype is always plural. One person’s excessive enthusiasm is a private bias; only when that enthusiasm becomes a shared emotional and cognitive state through mechanisms of social transmission—media, social networks, conferences, investor roadshows—does it become hype. Durkheim’s &lt;em&gt;collective effervescence&lt;/em&gt; applies here with some discomfort: at its peak, hype is phenomenologically almost isomorphic with the collective ecstasy of religious ritual.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Second, “overproduction.”&lt;/strong&gt; The metaphor is economic. Just as Marx located the core of capitalist overproduction crises in a structural imbalance between production and effective demand, the overproduction in hype is a spillover of expectations relative to current realizability. The qualifier &lt;em&gt;current&lt;/em&gt; is crucial: many hyped technologies did eventually deliver their wild early promises, only on a badly compressed timeline. The online shopping, video calls, and instant access to information imagined during the dot-com bubble all exist today—delayed by ten to fifteen years.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Third, “narrative construction.”&lt;/strong&gt; Hype never travels as a table of data. It travels as a story. Robert Shiller’s &lt;em&gt;Narrative Economics&lt;/em&gt; (2019) analyses this well: economic fluctuations are often driven by contagious narratives whose transmission dynamics resemble epidemiological models rather than the uniform diffusion of information assumed by rational expectations theory. A narrative that “AI will replace all white-collar work within five years” propagates far faster than a factual statement about where current LLMs sit on a particular benchmark.&lt;/p&gt;

&lt;h3 id=&quot;22-distinguishing-hype-from-its-neighbours&quot;&gt;2.2 Distinguishing hype from its neighbours&lt;/h3&gt;

&lt;p&gt;To sharpen the concept, it helps to separate hype from three things it is often confused with.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Hype ≠ optimism.&lt;/strong&gt; Optimism is a disposition toward positive expectation about the future; it may be grounded or ungrounded. Hype specifically denotes collective expectation that is amplified and accelerated by social transmission, appreciating in value as it spreads. An engineer’s optimism grounded in deep technical understanding and the excitement of someone who has never read a paper but is surrounded on Twitter by an “AGI is imminent” narrative are epistemologically quite different in texture, even if their outward behaviour—buying AI-related stocks, say—is identical.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Hype ≠ bubble.&lt;/strong&gt; A bubble is an economic concept: the systematic departure of asset prices from fundamental value. Hype can generate bubbles, but its extension is far wider. Cultural hype, technological hype, and political hype can all exist without any financial market involvement. The “phenomenal anticipation” around a film’s release is hype but involves no bubble. Conversely, participants in certain financial bubbles—the CDO market of 2007, for instance—were often aware that it was a bubble but unwilling to leave “before the music stopped,” which is closer to a coordination failure in game theory than to hype in the classical sense.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Hype ≠ propaganda.&lt;/strong&gt; Propaganda has a definite agent and a definite intent. What is uncanny about hype is that it is typically a subjectless process. Of course many hypes contain conscious boosters—founders, media, key opinion leaders—but the dynamics as a whole cannot be reduced to the intent of any single actor. As the inverse of the invisible hand, hype is an invisible mouth: everyone is only saying what they take to be true or interesting, and these local utterances aggregate into a self-reinforcing torrent of narrative whose direction and force are under no individual’s control.&lt;/p&gt;

&lt;h3 id=&quot;23-the-ontological-status-of-hype&quot;&gt;2.3 The ontological status of hype&lt;/h3&gt;

&lt;p&gt;Which raises a more fundamental philosophical question: is hype &lt;em&gt;real&lt;/em&gt;?&lt;/p&gt;

&lt;p&gt;In an important sense, yes. Hype may rest on a misjudgement of the future, but as a social-psychological phenomenon it is entirely real, and it produces entirely real consequences. The Thomas theorem applies: &lt;em&gt;if men define situations as real, they are real in their consequences.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;This is most conspicuous in finance. George Soros’s theory of reflexivity provides the most precise framework. In Soros’s model there is two-way causation between participants’ cognitive function and the participating function through which they act on the situation. When enough people believe a technology will change the world, capital flows in, talent clusters, infrastructure gets built—and all of that raises the probability that the technology &lt;em&gt;really will&lt;/em&gt; change the world. Hype is not merely a (possibly mistaken) prediction about the future; it is a force that alters the future.&lt;/p&gt;

&lt;p&gt;Which yields hype’s deepest paradox: &lt;mark&gt;a “false” collective belief can, by mobilising enough resources, make itself “true.”&lt;/mark&gt;&lt;/p&gt;

&lt;h2 id=&quot;iii-the-dynamics-of-hype-a-machine-that-feeds-itself&quot;&gt;III. The dynamics of hype: a machine that feeds itself&lt;/h2&gt;

&lt;h3 id=&quot;31-the-gartner-curve-and-its-limits&quot;&gt;3.1 The Gartner curve and its limits&lt;/h3&gt;

&lt;p&gt;Any discussion of hype dynamics has to pass through Gartner’s Hype Cycle, proposed in 1995, which divides a technology’s life into five phases: Innovation Trigger, Peak of Inflated Expectations, Trough of Disillusionment, Slope of Enlightenment, Plateau of Productivity.&lt;/p&gt;

&lt;figure&gt;
  &lt;svg viewBox=&quot;0 0 760 330&quot; width=&quot;760&quot; height=&quot;330&quot; role=&quot;img&quot; aria-labelledby=&quot;gartner-title-en&quot; style=&quot;max-width:100%;height:auto;display:block;font-family:var(--sans)&quot;&gt;
    &lt;title id=&quot;gartner-title-en&quot;&gt;The Gartner Hype Cycle: expectations inflate sharply to a peak, fall into a trough of disillusionment, then climb a slope of enlightenment to a plateau of productivity&lt;/title&gt;
    &lt;line x1=&quot;58&quot; y1=&quot;284&quot; x2=&quot;736&quot; y2=&quot;284&quot; stroke=&quot;var(--rule)&quot; stroke-width=&quot;1&quot;&gt;&lt;/line&gt;
    &lt;line x1=&quot;58&quot; y1=&quot;284&quot; x2=&quot;58&quot; y2=&quot;26&quot; stroke=&quot;var(--rule)&quot; stroke-width=&quot;1&quot;&gt;&lt;/line&gt;
    &lt;text x=&quot;58&quot; y=&quot;18&quot; fill=&quot;var(--ink-soft)&quot; font-size=&quot;11&quot; letter-spacing=&quot;0.08em&quot;&gt;EXPECTATIONS&lt;/text&gt;
    &lt;text x=&quot;736&quot; y=&quot;308&quot; fill=&quot;var(--ink-soft)&quot; font-size=&quot;11&quot; letter-spacing=&quot;0.08em&quot; text-anchor=&quot;end&quot;&gt;TIME&lt;/text&gt;
    &lt;path d=&quot;M 62 278 C 120 272 168 232 206 62 C 232 168 250 236 300 258 C 340 274 372 272 412 258 C 486 232 560 190 640 176 L 732 172&quot; fill=&quot;none&quot; stroke=&quot;var(--accent)&quot; stroke-width=&quot;2.5&quot; stroke-linecap=&quot;round&quot;&gt;&lt;/path&gt;
    &lt;circle cx=&quot;206&quot; cy=&quot;62&quot; r=&quot;4&quot; fill=&quot;var(--accent)&quot;&gt;&lt;/circle&gt;
    &lt;circle cx=&quot;330&quot; cy=&quot;264&quot; r=&quot;4&quot; fill=&quot;var(--accent)&quot;&gt;&lt;/circle&gt;
    &lt;text x=&quot;66&quot; y=&quot;252&quot; fill=&quot;var(--ink-soft)&quot; font-size=&quot;12&quot;&gt;Innovation trigger&lt;/text&gt;
    &lt;text x=&quot;206&quot; y=&quot;46&quot; fill=&quot;var(--ink-soft)&quot; font-size=&quot;12&quot; text-anchor=&quot;middle&quot;&gt;Peak of inflated expectations&lt;/text&gt;
    &lt;text x=&quot;330&quot; y=&quot;282&quot; fill=&quot;var(--ink-soft)&quot; font-size=&quot;12&quot; text-anchor=&quot;middle&quot;&gt;Trough of disillusionment&lt;/text&gt;
    &lt;text x=&quot;510&quot; y=&quot;238&quot; fill=&quot;var(--ink-soft)&quot; font-size=&quot;12&quot; text-anchor=&quot;middle&quot;&gt;Slope of enlightenment&lt;/text&gt;
    &lt;text x=&quot;732&quot; y=&quot;158&quot; fill=&quot;var(--ink-soft)&quot; font-size=&quot;12&quot; text-anchor=&quot;end&quot;&gt;Plateau of productivity&lt;/text&gt;
  &lt;/svg&gt;
  &lt;figcaption&gt;
    &lt;p&gt;The five phases of the Gartner Hype Cycle (1995). Schematic: the vertical axis is collective expectation, not any measurable property of the technology—which is precisely the problem discussed below.&amp;lt;/figcaption&amp;gt;
&amp;lt;/figure&amp;gt;&lt;/p&gt;

    &lt;p&gt;The model’s virtue is its intuitive plausibility: almost everyone can recognise the shape from experience. Blockchain, VR, the metaverse, 3D printing, nanotechnology—every technological concept seems to find its place somewhere on the curve.&lt;/p&gt;

    &lt;p&gt;As an analytical instrument, however, it has several fundamental defects.&lt;/p&gt;

    &lt;p&gt;&lt;strong&gt;First, it is descriptive and lacks explanatory power.&lt;/strong&gt; It tells you the shape of the curve without telling you why the curve has that shape. Why do expectations first overshoot and then undershoot? What are the underlying micro-mechanisms? On this the model is silent.&lt;/p&gt;

    &lt;p&gt;&lt;strong&gt;Second, it assumes a single-peaked, deterministic trajectory.&lt;/strong&gt; Reality is more complicated. Many technologies go through several hype cycles—AI alone has had at least three, in the 1960s, the 1980s, and the 2010s—each with a different shape, amplitude, and frequency. Others were forgotten entirely before reaching any plateau.&lt;/p&gt;

    &lt;p&gt;&lt;strong&gt;Third, and most important, it smuggles in a teleological assumption:&lt;/strong&gt; as though every technology were &lt;em&gt;destined&lt;/em&gt; to pass through all five phases and arrive at a rational plateau. This narrative quietly defines hype as an irrational stage that must be overcome, an unfortunate detour on the way to a rational terminus. But what if hype is an intrinsic component of how the system runs?&lt;/p&gt;

    &lt;h3 id=&quot;32-micro-mechanisms-information-cascades-and-social-proof&quot;&gt;3.2 Micro-mechanisms: information cascades and social proof&lt;/h3&gt;

    &lt;p&gt;To supply the explanatory power Gartner lacks, we need to descend to the micro-mechanisms of hype formation.&lt;/p&gt;

    &lt;p&gt;The 1992 paper on informational cascades by Sushil Bikhchandani, David Hirshleifer, and Ivo Welch provides the first piece. Their model shows that in a sequential decision environment, even if every individual is Bayesian-rational, it takes only a few early decision-makers happening to choose in the same direction for later decision-makers to rationally ignore their own private signals and follow their predecessors. Once such a cascade forms, it produces extreme informational fragility: the behaviour of an entire population may rest on a very small quantity of initial information.&lt;/p&gt;

    &lt;p&gt;Map this onto hype. When a handful of influential early voices—top-tier venture funds, well-known academics, technology media—express strong optimism about a technology, subsequent observers face an asymmetry. They cannot tell whether those early voices rest on deep technical assessment or on limited information plus cognitive bias. But under the logic of social proof, “so many smart people believe in it” is itself a powerful signal.&lt;/p&gt;

    &lt;p&gt;Robert Cialdini treats social proof as one of the most powerful decision heuristics available to us. In domains of high uncertainty—and the future value of a new technology is a paradigm case—reliance on social proof rises sharply. The result is a positive feedback loop:&lt;/p&gt;

    &lt;blockquote&gt;
      &lt;p&gt;the higher the uncertainty → the greater the reliance on others’ judgement → the more easily consensus forms → the more easily the consensus that forms departs from the true value.&lt;/p&gt;
    &lt;/blockquote&gt;

    &lt;h3 id=&quot;33-viral-narrative-transmission&quot;&gt;3.3 Viral narrative transmission&lt;/h3&gt;

    &lt;p&gt;Cascades explain why consensus forms easily, but not the characteristic &lt;em&gt;overheating&lt;/em&gt; of hype. For that we need the dynamics of narrative transmission.&lt;/p&gt;

    &lt;p&gt;Shiller borrows the epidemiological SIR model to analyse the spread of economic narratives. Transmission depends on three parameters: the contagion rate (the probability an “infected” individual passes the narrative on), the recovery rate (the probability an individual stops transmitting it), and the initial infected base.&lt;/p&gt;

    &lt;p&gt;The key insight is that contagion rates differ. Narratives with the following features spread fastest.&lt;/p&gt;

    &lt;ol&gt;
  &lt;li&gt;&lt;strong&gt;High emotional arousal.&lt;/strong&gt; “AI could put three hundred million people out of work within a decade” provokes far more fear and excitement than “AI has reached human-level accuracy on a specific task.”&lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Identity provision.&lt;/strong&gt; “Are you someone who understands the future, or someone the future will discard?” Binary narratives of this kind bind a person’s technological position to their self-conception, which raises stickiness enormously.&lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Simplicity and retellability.&lt;/strong&gt; “The metaverse is the next internet” travels across a dinner table far more easily than any technical white paper. And each retelling simplifies, dramatises, and sharpens it—what folklorists studying serial reproduction have long described.&lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Time pressure.&lt;/strong&gt; “The window is short.” “First-mover advantage is decisive.” “If you don’t move now it will be too late.” Scarcity narratives exploit loss aversion.&lt;/li&gt;
&lt;/ol&gt;

    &lt;p&gt;When a narrative has all four at once, its contagion rate becomes very high while its recovery rate is pressed very low—because exiting the narrative means losing an identity and incurring the fear of missing out. This is the overheating mechanism: positive feedback in narrative transmission inflates collective expectations far faster than information actually accumulates.&lt;/p&gt;

    &lt;h3 id=&quot;34-feedback-loops-and-self-fulfilment&quot;&gt;3.4 Feedback loops and self-fulfilment&lt;/h3&gt;

    &lt;p&gt;There is a further layer. As Soros’s reflexivity implies, hype does not merely reflect reality; it alters it.&lt;/p&gt;

    &lt;p&gt;Take AI. Once the “AI will change everything” narrative reaches sufficient intensity, the following begin to happen:&lt;/p&gt;

    &lt;ul&gt;
  &lt;li&gt;capital pours in, so more research gets funded and more companies get financed;&lt;/li&gt;
  &lt;li&gt;talent flows toward AI, accelerating the actual rate of technical progress;&lt;/li&gt;
  &lt;li&gt;governments come to see AI as a focus of strategic competition and issue supportive policy;&lt;/li&gt;
  &lt;li&gt;universities expand enrolment in AI-related programmes;&lt;/li&gt;
  &lt;li&gt;public acceptance of AI products rises, lowering the cost of market entry.&lt;/li&gt;
&lt;/ul&gt;

    &lt;p&gt;These real changes then supply the hype narrative with fresh “evidence”: look, the capital is flowing in; look, the best people are switching fields; look, even governments are paying attention. A self-fulfilling prophecy takes shape—though the degree and direction of that fulfilment may diverge widely from the original narrative.&lt;/p&gt;

    &lt;p&gt;Which is hype’s most unsettling epistemological predicament: &lt;mark&gt;while you are inside the rising phase, all the evidence points to the reasonableness of the hype, because the hype is manufacturing the evidence.&lt;/mark&gt; You cannot judge whether hype is excessive by observing current evidence, because you cannot separate “capital flowed in because the technology is genuinely good” from “capital flowed in because the narrative said the technology is good—and the inflow then made the technology look better.”&lt;/p&gt;

    &lt;h3 id=&quot;35-what-triggers-collapse&quot;&gt;3.5 What triggers collapse&lt;/h3&gt;

    &lt;p&gt;If the positive feedback loop is so powerful, why does it ever collapse?&lt;/p&gt;

    &lt;p&gt;Because the stability of a positive feedback system depends on a continuing inflow of new “infections.” Once the susceptible population is exhausted—everyone persuadable by the narrative has been persuaded—the increment of transmission begins to fall. Meanwhile the negative feedback has never gone away; it has only been suppressed:&lt;/p&gt;

    &lt;ul&gt;
  &lt;li&gt;&lt;strong&gt;Accumulated reality gaps.&lt;/strong&gt; Early promises come due. Products fall short. The ROI does not materialise. The use cases do not appear. Every gap is a small seed of counter-narrative.&lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Defectors.&lt;/strong&gt; Some early participants begin to voice doubt in public, and their defection carries asymmetric weight: a former believer saying “I was wrong” is more contagious than a lifelong sceptic saying “I told you so.”&lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Competition for attention.&lt;/strong&gt; Collective attention is finite. A new hype topic drains energy from the old one.&lt;/li&gt;
&lt;/ul&gt;

    &lt;p&gt;When the increment of positive feedback falls below that of negative feedback, the system reverses. And because the same positive-feedback machinery now runs backwards—“the smartest people are leaving,” “investors are pulling out,” “the press has turned”—the descent is usually faster than the ascent. That is why the trough of disillusionment on the Gartner curve tends to be steeper than the peak.&lt;/p&gt;

    &lt;h2 id=&quot;iv-the-political-economy-of-hype-who-produces-it-who-consumes-it&quot;&gt;IV. The political economy of hype: who produces it, who consumes it?&lt;/h2&gt;

    &lt;h3 id=&quot;41-the-production-side&quot;&gt;4.1 The production side&lt;/h3&gt;

    &lt;p&gt;The emphasis above on hype as a subjectless process does not mean there is no structural distribution of interests inside it. Hype has a clear political economy.&lt;/p&gt;

    &lt;p&gt;&lt;strong&gt;Venture capital is the most important institutionalised producer of modern hype.&lt;/strong&gt; The VC business model depends intrinsically on producing it. A power-law return structure means a fund needs a few investments to return enormously in order to cover many failures. To return enormously, a portfolio company must raise later rounds at higher valuations and eventually exit through IPO or acquisition. And higher valuations depend heavily on the market’s expectations about the future growth of that company’s category.&lt;/p&gt;

    &lt;p&gt;In other words, VCs have a structural incentive to produce and maintain hype. This requires no individual VC to be insincere: one can sincerely believe the category will change the world while objectively functioning as a transmission node for hype. &lt;mark&gt;There is no contradiction between institutional logic and personal sincerity.&lt;/mark&gt;&lt;/p&gt;

    &lt;p&gt;&lt;strong&gt;Technology media is the second great producer, and its incentives deserve the same dissection.&lt;/strong&gt; Media business models depend on capturing attention, and in the attention economy the mild narrative—“this technology may develop somewhat over the next decade”—never beats the violent one—“this technology is about to upend everything.” The result is a systematic selection bias toward amplifying the most extreme voices, because extreme voices generate the most clicks, shares, and engagement.&lt;/p&gt;

    &lt;p&gt;Note that this bias is &lt;strong&gt;symmetric&lt;/strong&gt;: media amplify the rising phase (“AI will change everything”) and the falling phase (“the AI bubble is about to burst”) alike. Their interest is not in sustaining any particular direction of narrative but in sustaining its extremity. Moderate middle positions do not generate traffic.&lt;/p&gt;

    &lt;p&gt;&lt;strong&gt;Academia plays an ambiguous role.&lt;/strong&gt; On one hand it positions itself as a rational producer of knowledge, immune to hype. On the other, grant funding requires researchers to argue for the transformative potential and broader impact of their work, which institutionally rewards systematic exaggeration of significance. Daniel Sarewitz (2016), in his critique of the science funding system, put it sharply: the competitive structure of modern research funding is in effect a machine for converting scientific knowledge into promissory narratives.&lt;/p&gt;

    &lt;h3 id=&quot;42-the-consumption-side&quot;&gt;4.2 The consumption side&lt;/h3&gt;

    &lt;p&gt;Consumers of hype have their own complex motivational structure.&lt;/p&gt;

    &lt;p&gt;&lt;strong&gt;For the general public&lt;/strong&gt;, hype supplies a cheap cognitive frame. In a world of information overload, genuinely understanding what a new technology means requires enormous cognitive investment. A hype narrative offers a shortcut: “you don’t need to understand the details of the transformer architecture, you just need to know that AI is about to change everything.” This cognitive offloading explains why hype spreads especially violently in highly specialised fields: the higher the barrier to expertise, the greater the dependence on simplified narrative.&lt;/p&gt;

    &lt;p&gt;&lt;strong&gt;For corporate decision-makers&lt;/strong&gt;, hype supplies legitimacy. Under high uncertainty, “everyone is doing X” is itself the best reason to do X. This is not entirely irrational: in a world with network effects and ecosystem lock-in, “right direction, late start” can be more fatal than “perfect judgement, missed window.” Which is why FOMO is hype’s most faithful co-pilot.&lt;/p&gt;

    &lt;p&gt;&lt;strong&gt;For individual practitioners&lt;/strong&gt;, hype supplies an identity narrative and a sense of career direction. “I work in AI” carries far more social identity value at the peak than “I am a database engineer,” even where the latter’s actual skills may be scarcer. This identity dimension makes exit especially hard: admitting that “the direction I have been chasing may not be as revolutionary as I thought” requires not merely a revision of opinion but a reconstruction of self-narrative.&lt;/p&gt;

    &lt;h3 id=&quot;43-redistributive-effects&quot;&gt;4.3 Redistributive effects&lt;/h3&gt;

    &lt;p&gt;A political economy of hype cannot avoid the basic question: who gains and who loses?&lt;/p&gt;

    &lt;p&gt;Broadly, hype produces a systematic redistribution of wealth and attention:&lt;/p&gt;

    &lt;ul&gt;
  &lt;li&gt;&lt;strong&gt;From late participants to early ones.&lt;/strong&gt; Most nakedly in finance: later investors enter at higher valuations and bear the largest losses if the hype collapses.&lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;From outsiders to insiders.&lt;/strong&gt; Those who can produce hype narratives—investors, founders, industry analysts—typically get more information and better exit options than those who consume them.&lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;From the neglected alternatives to the hyped direction.&lt;/strong&gt; This is the most hidden and possibly the most important cost: when resources pour into the hyped direction, fields that were never hyped but may be equally or more important go hungry. It is an invisible opportunity cost.&lt;/li&gt;
&lt;/ul&gt;

    &lt;h2 id=&quot;v-the-deep-historical-structure-of-hype&quot;&gt;V. The deep historical structure of hype&lt;/h2&gt;

    &lt;h3 id=&quot;51-hype-is-not-a-modern-phenomenon&quot;&gt;5.1 Hype is not a modern phenomenon&lt;/h3&gt;

    &lt;p&gt;If hype is the overproduction of collective expectations, its history long predates Silicon Valley and venture capital.&lt;/p&gt;

    &lt;figure&gt;
      &lt;p&gt;&lt;img src=&quot;/assets/images/brueghel-tulip-mania.jpg&quot; alt=&quot;Oil painting: monkeys in seventeenth-century Dutch dress trade tulips in a garden—some poring over ledgers, some counting coins, one hauled before a court, one urinating on a worthless bulb&quot; width=&quot;1200&quot; height=&quot;767&quot; /&gt;&lt;/p&gt;
      &lt;figcaption&gt;
        &lt;p&gt;Jan Brueghel the Younger, &lt;em&gt;Satire on Tulip Mania&lt;/em&gt;, c. 1640. Oil on panel, Frans Hals Museum, Haarlem. Public domain, via Wikimedia Commons. Every speculator in the picture is a monkey in human clothes—the earliest known painting made specifically to mock a hype.&amp;lt;/figcaption&amp;gt;
&amp;lt;/figure&amp;gt;&lt;/p&gt;

        &lt;p&gt;Dutch tulipmania (1637) is usually treated as the first well-documented speculative bubble. But Anne Goldgar’s revisionist history &lt;em&gt;Tulipmania&lt;/em&gt; (2007) argues that the traditional account was itself enormously hyped by posterity: far fewer people actually speculated than later literature implies, and the price explosion was concentrated in a handful of rare varieties. &lt;mark&gt;A story about hype was itself hyped.&lt;/mark&gt; This meta-level irony is nearly standard equipment for the phenomenon.&lt;/p&gt;

        &lt;p&gt;The South Sea Bubble (1720) offers a cleaner dissection. The company’s core narrative was that it held a monopoly on trade with South America—then regarded as a new world of inexhaustible riches—which would generate unlimited profit. That narrative had every high-contagion feature listed above: extreme emotional arousal (overnight wealth), identity provision (becoming a visionary of the new age), simplicity (“South America has endless gold and silver”), and time pressure (“the stock rises daily; buy today or pay more tomorrow”).&lt;/p&gt;

        &lt;p&gt;Newton’s famous loss in the South Sea Bubble—reportedly around £20,000, equivalent to several million today—is often cited to show that even the cleverest are fooled by hype. But the deeper lesson may be this: Newton’s genius equipped him to analyse deterministic systems, and hype is a reflexive social system. The tools that work on the former fail precisely here.&lt;/p&gt;

        &lt;h3 id=&quot;52-why-modernity-intensifies-hype&quot;&gt;5.2 Why modernity intensifies hype&lt;/h3&gt;

        &lt;p&gt;If hype is a structural feature of human cognition, why does contemporary hype feel more violent and more frequent than at any point in history?&lt;/p&gt;

        &lt;p&gt;&lt;strong&gt;The exponential acceleration of transmission.&lt;/strong&gt; From the mail coach to the telegraph to radio to television to the internet to social media, the latency of information has been compressed to near real time, which sharply accelerates each iteration of the positive feedback loop. Tulipmania took years to peak; a crypto cycle can complete in months or even weeks.&lt;/p&gt;

        &lt;p&gt;&lt;strong&gt;Deepening financialization.&lt;/strong&gt; The modern financial system has created ever more instruments for trading on expectations. Options, futures, leveraged ETFs, crypto tokens—each is a mechanism for converting expectation into a position, and the profit or loss on that position feeds back to reinforce or dissolve the expectation. The more financialized the system, the stronger the reflexive loop.&lt;/p&gt;

        &lt;p&gt;&lt;strong&gt;The extremity of specialisation and the cognitive division of labour.&lt;/strong&gt; In a highly specialised society, no individual can independently evaluate knowledge claims across all fields. Dependence on “expert opinion” and “social consensus” is therefore higher than ever—which is exactly the breeding ground for cascades and for hype.&lt;/p&gt;

        &lt;p&gt;&lt;strong&gt;The democratisation (and industrialisation) of narrative production.&lt;/strong&gt; Social media gave everyone the ability to produce and transmit narrative, and simultaneously created a large professional class whose occupation &lt;em&gt;is&lt;/em&gt; narrative production—opinion leaders, content creators, industry analysts. The supply side expanded enormously while the quality-control mechanisms—editorial review in traditional media, peer review in academia—were sharply weakened.&lt;/p&gt;

        &lt;h2 id=&quot;vi-the-epistemology-of-hype-can-we-get-beyond-it&quot;&gt;VI. The epistemology of hype: can we get beyond it?&lt;/h2&gt;

        &lt;h3 id=&quot;61-the-counter-hype-trap&quot;&gt;6.1 The counter-hype trap&lt;/h3&gt;

        &lt;p&gt;The natural response to hype is counter-hype: stand opposite it and stay coolly sceptical. But counter-hype has epistemological traps of its own.&lt;/p&gt;

        &lt;p&gt;&lt;strong&gt;Counter-hype may be hype in reverse.&lt;/strong&gt; When someone’s identity rests on being “the one who saw through the hype,” he acquires a systematic incentive to deny every optimistic signal. “AI is just statistics.” “Crypto is just a Ponzi scheme.” “EVs are just toys.” These are extreme simplifications of a complex reality, exactly like “AI will change everything,” “crypto will upend finance,” “EVs will kill the combustion engine.” &lt;mark&gt;Scepticism enjoys no epistemological privilege.&lt;/mark&gt;&lt;/p&gt;

        &lt;p&gt;Nassim Taleb’s intellectual legacy makes an interesting case study here. His critique of the narrative fallacy and of overconfidence is genuinely penetrating, yet his writing style and public persona rest on a hype about being anti-hype: a self-narrative about being more clear-eyed than everyone else, whose transmission dynamics are nearly isomorphic with what it criticises.&lt;/p&gt;

        &lt;p&gt;&lt;strong&gt;Counter-hype makes errors of timescale.&lt;/strong&gt; As noted, many hyped technologies did deliver on their early promises, only over a far longer horizon than the hype assumed. Roy Amara’s law captures it exactly: we tend to overestimate the effect of a technology in the short run and underestimate it in the long run. In this sense, “AI is over-hyped” may be right in the short run and wrong in the long run, while “AI will change everything” may be right in the long run and dangerous in the short. Correct analysis requires precise calibration of timescale—which happens to be one of the weakest points in human cognition.&lt;/p&gt;

        &lt;h3 id=&quot;62-a-more-mature-epistemological-posture&quot;&gt;6.2 A more mature epistemological posture&lt;/h3&gt;

        &lt;p&gt;If neither embracing hype blindly nor opposing it wholesale is a good strategy, what is?&lt;/p&gt;

        &lt;p&gt;I favour a position that might be called &lt;strong&gt;participatory skepticism&lt;/strong&gt;, with five elements.&lt;/p&gt;

        &lt;p&gt;&lt;strong&gt;First, acknowledge your own embeddedness.&lt;/strong&gt; You are not outside the hype. Your sources, your social circle, your professional interests, your identity—all place you at a particular position in the field. Acknowledging that embeddedness is a more honest and more useful starting point than pretending to a view from nowhere.&lt;/p&gt;

        &lt;p&gt;&lt;strong&gt;Second, separate directional judgements from temporal ones.&lt;/strong&gt; “Does this technological direction have value?” and “on what timescale will that value be realised?” are independent questions, and hype almost always conflates them. A technology can be simultaneously right in the long run and badly over-hyped in the short. Holding both requires cognitive flexibility.&lt;/p&gt;

        &lt;p&gt;&lt;strong&gt;Third, watch changes in material infrastructure more closely than changes in narrative.&lt;/strong&gt; Narrative can flip overnight; material infrastructure—laboratories, factories, talent pipelines, regulatory frameworks, user habits—changes slowly and stickily. When you are trying to find signal in the noise, the most reliable indicators are usually these slow variables. Is real R&amp;amp;D spending rising or falling? Are key people flowing in or out? Is user retention rising or falling?&lt;/p&gt;

        &lt;p&gt;&lt;strong&gt;Fourth, think in probability distributions and resist the pull of binary narrative.&lt;/strong&gt; “Either AI changes everything or it is just a bubble” is a false dichotomy. The truer picture is a distribution: different degrees of impact, in different domains, on different timescales, each with its own likelihood. Philip Tetlock’s work in &lt;em&gt;Superforecasting&lt;/em&gt; (2015) shows that the best forecasters are precisely those most given to fine-grained probability estimates, least drawn to extreme positions, and most willing to update.&lt;/p&gt;

        &lt;p&gt;&lt;strong&gt;Fifth, keep metacognitive watch on your own degree of certainty.&lt;/strong&gt; When you feel very sure that something is “definitely hype” or “definitely going to happen,” that feeling is itself a signal worth auditing. The intensity of certainty has no reliable correlation with the accuracy of judgement. Dunning–Kruger applies not only to assessments of competence but to assessments of hype: the more confident someone is about judging hype, the less they usually understand its complexity.&lt;/p&gt;

        &lt;h2 id=&quot;vii-the-existential-dimension-why-do-people-need-hype&quot;&gt;VII. The existential dimension: why do people need hype?&lt;/h2&gt;

        &lt;h3 id=&quot;71-hype-as-a-secular-religion&quot;&gt;7.1 Hype as a secular religion&lt;/h3&gt;

        &lt;p&gt;Raise the level of analysis one more step—out of sociology and economics into anthropology and ontology—and an uncomfortable but possibly important hypothesis emerges: &lt;mark&gt;hype satisfies something like a religious need.&lt;/mark&gt;&lt;/p&gt;

        &lt;p&gt;Peter Thiel returns repeatedly in his public talks to a theme: the fundamental problem of the modern world is the exhaustion of imagination about the future. He argues that mid-twentieth-century Western society possessed a strong &lt;em&gt;definite optimism&lt;/em&gt;—a belief in a concrete, plannable good future (the moon programme, the interstate highway system, the peaceful use of nuclear energy). That belief dissolved after the 1970s and was replaced by an &lt;em&gt;indefinite optimism&lt;/em&gt;: the future will be better, but nobody knows how.&lt;/p&gt;

        &lt;p&gt;If part of Thiel’s diagnosis is right, hype can be understood as a compensatory response to the scarcity of definite optimism. In a “postmodern” world where the grand narrative has died (Lyotard, 1979), every round of technological hype is a brief resurrection of one. “AI will change everything.” “Blockchain will rebuild trust.” “The metaverse will surpass reality.” These narratives supply a sense of direction, of meaning, of a future one can get hold of—precisely what secular modernity finds it increasingly hard to provide.&lt;/p&gt;

        &lt;p&gt;Seen this way, the recurrence of hype cycles—including the fact that a new hype topic always arrives after each disillusionment—is not merely a product of market dynamics. It reflects a deep human hunger for a collective sense of purpose. That hunger is structural and ineliminable, and so, therefore, is hype.&lt;/p&gt;

        &lt;h3 id=&quot;72-hype-and-temporality&quot;&gt;7.2 Hype and temporality&lt;/h3&gt;

        &lt;p&gt;Heidegger’s analysis of human temporality offers another angle. For Heidegger, human existence is essentially directed toward the future (&lt;em&gt;Sein-zum-Tode&lt;/em&gt;, being-toward-death). We are always already projecting (&lt;em&gt;Entwurf&lt;/em&gt;); our present actions, moods, and judgements are permeated by anticipation of what is to come.&lt;/p&gt;

        &lt;p&gt;Hype can be read as a collective expression of that ontological future-directedness. Human beings do not merely happen to be prone to excessive expectation about the future; in a sense, expectant projection toward the future &lt;em&gt;is&lt;/em&gt; a basic structure of human existence. Hype is that structure developing, like a photograph, under particular social and historical conditions.&lt;/p&gt;

        &lt;p&gt;This does not mean we should embrace hype or abandon criticism. It means that &lt;strong&gt;criticism of hype cannot stop at “people are too easily fooled.”&lt;/strong&gt; Behind the being-fooled is a fundamental need for meaning, direction, and hope, and that need does not disappear because you have pointed out the irrationality. Any project to eliminate hype that cannot supply an alternative mechanism for producing meaning is bound to fail.&lt;/p&gt;

        &lt;h2 id=&quot;viii-coda-living-with-hype&quot;&gt;VIII. Coda: living with hype&lt;/h2&gt;

        &lt;p&gt;Return to where we began: the identification of hype may itself be another form of hype.&lt;/p&gt;

        &lt;p&gt;Having written this far, I have to admit a self-reflexive predicament. Is this essay not also a hype? A narrative about &lt;em&gt;deeply understanding hype&lt;/em&gt;, offering in academic wrapping a feeling of cognitive superiority, so the reader may feel clearer-eyed than all those people swept along by it?&lt;/p&gt;

        &lt;p&gt;Yes. To some degree, yes.&lt;/p&gt;

        &lt;p&gt;But admitting it does not invalidate the analysis. It only reminds us that &lt;mark&gt;there is no Archimedean point with respect to hype. We are permanently embedded in the system we are trying to analyse. This is not a defect of cognition; it is the condition of cognition.&lt;/mark&gt;&lt;/p&gt;

        &lt;p&gt;A mature intellectual posture might look like this.&lt;/p&gt;

        &lt;p&gt;You know how hype works. You know about information cascades, social proof, narrative transmission, reflexive loops. You know your own judgement is contaminated by your social position and your identity. You know that being anti-hype can be as wrong as embracing it. You know that all of this knowing may not help you judge any better.&lt;/p&gt;

        &lt;p&gt;And then you still have to act under uncertainty.&lt;/p&gt;

        &lt;p&gt;You still have to make decisions about direction, timing, and the commitment of resources, with incomplete information, with signal hard to separate from noise, and with a future that is not in principle predictable. All you can do is stay alert to the limits of your own cognition, distrust the feeling of certainty, and keep a respect for complexity.&lt;/p&gt;

        &lt;p&gt;To adapt the famous prayer from Reinhold Niebuhr:&lt;/p&gt;

        &lt;blockquote&gt;
          &lt;p&gt;Grant me the courage to join the hypes worth joining, the wisdom to avoid the ones that are collapsing, and the insight to know the difference.&lt;/p&gt;
        &lt;/blockquote&gt;

        &lt;p&gt;Insight, of course, may also be over-hyped.&lt;/p&gt;
      &lt;/figcaption&gt;
    &lt;/figure&gt;
  &lt;/figcaption&gt;
&lt;/figure&gt;
</content>
  </entry>
  
  <entry>
    <title>我们究竟如何知道一个模型「更强」了？</title>
    <link href="https://mochiaochen.github.io/writing/2026/08/how-we-know-a-model-is-better/" rel="alternate" type="text/html"/>
    <published>2026-08-29T10:00:00+08:00</published>
    <updated>2026-08-29T10:00:00+08:00</updated>
    <id>https://mochiaochen.github.io/writing/2026/08/how-we-know-a-model-is-better</id>
    <content type="html" xml:base="https://mochiaochen.github.io/writing/2026/08/how-we-know-a-model-is-better/">&lt;p&gt;如果把过去几年的大模型竞赛压缩成几个问题，大概会经历这样的变化：最早大家问的是 &lt;em&gt;How big is the model?&lt;/em&gt;，后来变成 &lt;em&gt;How much compute did you use?&lt;/em&gt;，再后来是 &lt;em&gt;What is your MMLU score?&lt;/em&gt;。而当模型真的开始进入 coding、research、finance、customer support，甚至直接操作电脑和浏览器之后，一个更麻烦的问题逐渐浮现出来：&lt;/p&gt;

&lt;blockquote&gt;
  &lt;p&gt;&lt;strong&gt;How do we actually know whether the model is getting better?&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;这就是 Evaluation，或者大家更习惯说的 Eval。&lt;/p&gt;

&lt;p&gt;Eval 表面上很好理解：拿一套题给模型做，然后算分。但如果真正参与过模型研发，你很快就会发现，Evaluation 可能是整个 LLM pipeline 里最容易被低估、同时又最接近「核心」的一环。因为它不只是训练完成以后拿来验收模型的 QA system。&lt;mark&gt;你怎么定义「好」，最终就决定了团队收什么数据、怎样做 post-training、reward model 学什么、RL 优化什么，甚至 inference time 应该搜索什么。&lt;/mark&gt;&lt;/p&gt;

&lt;p&gt;换句话说：&lt;/p&gt;

&lt;blockquote&gt;
  &lt;p&gt;&lt;strong&gt;The model you get is downstream of the eval you build.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;从这个意义上说，大模型研发真正的起点，并不一定是 training，而是 measurement。&lt;/p&gt;

&lt;h2 id=&quot;01eval-不是-benchmark先搞清楚我们到底在测什么&quot;&gt;01｜Eval 不是 Benchmark：先搞清楚我们到底在测什么&lt;/h2&gt;

&lt;p&gt;很多人第一次接触 Evaluation，会把 benchmark、test set、leaderboard、eval 几个词混在一起。实际上 benchmark 只是 eval 的一个组成部分。一个完整的 Evaluation System，至少可以写成一个六元组：&lt;/p&gt;

&lt;p&gt;[\text{Eval} \;=\; \big\langle\, C,\; T,\; G,\; J,\; M,\; P \,\big\rangle]&lt;/p&gt;

&lt;p&gt;其中 (C) 是想测量的 construct（能力本身），(T) 是代表这个能力的 task distribution，(G) 是 ground truth 或 rubric，(J) 是做判断的 judge，(M) 是把判断聚合成数字的 metric，(P) 则是模型接受测试时的 protocol（prompt 模板、few-shot 数量、temperature、工具权限、上下文长度）。&lt;/p&gt;

&lt;p&gt;也就是说，你至少需要回答六个问题：想测什么能力？用哪些任务代表这个能力？什么样的回答叫好？谁来判断？怎么把判断变成数字？模型在什么条件下接受测试？&lt;/p&gt;

&lt;p&gt;以最经典的 MMLU 为例，它把模型放进一个高度标准化的考试场景：57 个 task，覆盖 elementary mathematics、US history、computer science、law 等不同学科，用 multiple-choice accuracy 测试模型在广泛知识和问题求解上的表现。MMLU 在 2020 年提出时，大模型在这些任务上距离 expert-level performance 还有非常大的差距，因此它可以很好地区分不同模型。&lt;/p&gt;

&lt;p&gt;但这里马上出现了 Evaluation 中最重要的一个概念：&lt;strong&gt;construct validity&lt;/strong&gt;，构念效度。&lt;/p&gt;

&lt;p&gt;MMLU 分数高，究竟说明了什么？它至少说明模型在这一组多学科选择题上表现不错。但它是不是意味着模型「更聪明」？是不是意味着它更适合当 research assistant？是不是更会 coding？是不是更擅长真实世界的 planning？这些结论都不能直接推出。&lt;/p&gt;

&lt;p&gt;这也是所有 Evaluation 的第一原则：&lt;/p&gt;

&lt;blockquote&gt;
  &lt;p&gt;&lt;strong&gt;Never confuse the metric with the construct.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;指标只是我们为了测量某个抽象能力而设计的 proxy，而不是能力本身。用符号写出来，我们真正关心的是 (C)，但能观测到的只有 (M)，而两者之间永远隔着一层：&lt;/p&gt;

&lt;p&gt;[M \;=\; f(C) + \varepsilon]&lt;/p&gt;

&lt;p&gt;其中 (\varepsilon) 包含了 task sampling 的偏差、grader 的噪声、prompt 格式的敏感度、以及数据污染。Eval engineering 的大部分工作，本质上就是在压缩这个 (\varepsilon)，并且诚实地承认它没有被压到零。&lt;/p&gt;

&lt;p&gt;比如大家经常说要测 reasoning。但 reasoning 本身至少可以继续拆成 deductive reasoning、mathematical reasoning、causal reasoning、multi-hop reasoning、planning、counterfactual reasoning、long-horizon reasoning。你如果连自己到底想测哪一种 reasoning 都没有定义清楚，那么后面的 dataset 再大、grader 再高级、统计方法再漂亮，其实都没有太大意义。&lt;/p&gt;

&lt;p&gt;所以真正成熟的 eval 从来不是从「找哪个 benchmark 跑一下」开始，而是从一句话开始：&lt;/p&gt;

&lt;blockquote&gt;
  &lt;p&gt;&lt;strong&gt;What capability or behavior are we trying to measure?&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2 id=&quot;02从-mmlu-到-humaneval为什么正确答案还不够&quot;&gt;02｜从 MMLU 到 HumanEval：为什么「正确答案」还不够？&lt;/h2&gt;

&lt;p&gt;MMLU 代表了非常经典的一类 benchmark：题目是静态的，答案是确定的，评分函数也很简单——Exact Match / Accuracy。&lt;/p&gt;

&lt;p&gt;这类 eval 有一个巨大的工程优势：它非常干净。只要协议一致，A 模型 80%，B 模型 85%，比较相对容易。&lt;/p&gt;

&lt;p&gt;但当模型从「回答问题」走向「创造一个可以执行的东西」以后，字符串匹配就开始失效了。&lt;/p&gt;

&lt;p&gt;2021 年 OpenAI 在 Codex 论文中提出 HumanEval。核心变化不是「把题目换成了编程题」，而是评分哲学变了：模型根据 docstring 生成 Python function，评估的重点是 &lt;strong&gt;functional correctness&lt;/strong&gt;——代码是不是真的能工作，而不是生成的文本像不像某个 reference answer。原论文中 Codex 在 HumanEval 上解决了 28.8% 的问题，而一个很有意思的结果是：如果允许同一个问题反复 sampling 100 次，至少找到一个可行解的比例可以提升到 70.2%。&lt;/p&gt;

&lt;p&gt;这件事非常重要，因为它实际上同时预示了两条后来越来越关键的路线。&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;第一条是：对于可以执行的任务，environment 本身就是最好的 evaluator。&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;你问「这段代码好不好」，让另一个 LLM 看一眼当然可以；但如果真正的问题是「这段代码能不能完成需求」，最可信的方法通常仍然是：&lt;em&gt;Run the tests.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;这就是 deterministic eval，包括 exact match、unit tests、compiler / runtime、SQL execution、schema validation、numerical tolerance、state checking。&lt;/p&gt;

&lt;p&gt;&lt;mark&gt;如果任务存在客观、可执行的 ground truth，优先使用 deterministic evaluator，往往比 LLM-as-a-Judge 更可靠。&lt;/mark&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;第二条则更加有意思：HumanEval 已经说明，模型能力不是一个单次 deterministic output，而是一个 probability distribution。&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;对于同一个问题 (x)，我们实际上面对的是&lt;/p&gt;

&lt;p&gt;[y \;\sim\; p_\theta(\,\cdot \mid x\,)]&lt;/p&gt;

&lt;p&gt;于是「第一次回答就正确的概率」可以写成&lt;/p&gt;

&lt;p&gt;[\text{pass@}1 \;=\; \mathbb{E}&lt;em&gt;{x}\Big[\ \mathbb{E}&lt;/em&gt;{y \sim p_\theta(\cdot\mid x)}\big[\ \mathbf{1}{\text{correct}(y)}\ \big]\Big]]&lt;/p&gt;

&lt;p&gt;而「给模型 (k) 次机会能不能找到正确答案」是另一回事。若对每题采样 (n) 个候选、其中 (c) 个正确，Codex 论文给出的无偏估计是&lt;/p&gt;

&lt;p&gt;[\text{pass@}k \;=\; \mathbb{E}_{x}\left[\, 1 - \frac{\dbinom{n-c}{k}}{\dbinom{n}{k}} \,\right]]&lt;/p&gt;

&lt;p&gt;这条线最终会一路发展到 Best-of-N、self-consistency、search、verifier、test-time compute。&lt;/p&gt;

&lt;p&gt;换句话说，从 HumanEval 开始，我们已经不再只评估：&lt;/p&gt;

&lt;blockquote&gt;
  &lt;p&gt;&lt;strong&gt;Can the model answer this question?&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;而开始评估：&lt;/p&gt;

&lt;blockquote&gt;
  &lt;p&gt;&lt;strong&gt;Can the system search its output space until it finds a correct answer?&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2 id=&quot;03从-humaneval-到-swe-bench会写代码不等于会做-software-engineering&quot;&gt;03｜从 HumanEval 到 SWE-bench：会写代码，不等于会做 Software Engineering&lt;/h2&gt;

&lt;p&gt;HumanEval 解决了一个重要问题：不要比较代码长得像不像 reference，应该直接测试 functional correctness。但它仍然有一个明显限制：它主要测试的是相对独立的函数生成任务。&lt;/p&gt;

&lt;p&gt;真实软件工程不是这样的。现实中的 programmer 接到的可能是：「这个 repository 有个 GitHub issue，用户说 pagination 在某种情况下出 bug，你去修一下。」&lt;/p&gt;

&lt;p&gt;于是模型需要先理解 issue，再阅读一个可能有几万行代码的 repo，定位 relevant files，理解现有 architecture，修改一个或多个文件，运行 tests，然后确保没有引入 regression。&lt;/p&gt;

&lt;p&gt;这就是 SWE-bench 想测的东西。原始 SWE-bench 从 12 个真实 Python repositories 中收集了 2,294 个真实 GitHub issues 及对应 pull requests。模型拿到的是 codebase 和 issue description，任务不是「写一个函数」，而是直接修改 repository，让 issue 被真正解决。这个过程经常要求跨 function、class 甚至多个 files 协同修改。论文最初发表时，即使当时最好的系统也只能解决很少一部分问题。&lt;/p&gt;

&lt;p&gt;这其实代表了整个 benchmark 设计思想的一次巨大迁移：&lt;/p&gt;

&lt;blockquote&gt;
  &lt;p&gt;HumanEval 测的是 &lt;strong&gt;code generation&lt;/strong&gt;；SWE-bench 测的是 &lt;strong&gt;software engineering task completion&lt;/strong&gt;。&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;而今天 SWE-bench Verified 又进一步加入了 human validation：从原始数据中筛出 500 个经过人工检查的实例，确保 issue description 清晰、test patch 合理，而且任务确实可以在所给信息下解决。&lt;/p&gt;

&lt;p&gt;这里出现了 Eval Engineering 中一个经常被忽略的问题：&lt;strong&gt;benchmark 本身也会有 bug&lt;/strong&gt;。&lt;/p&gt;

&lt;p&gt;如果一道题其实无法完成、reference answer 有问题、grader 写错了，那么你最后测到的就不是模型能力，而是 dataset noise。因此高质量 eval 并不是简单地「收更多题」，而是要不断做 dataset validation、error analysis、human audit、versioning。&lt;/p&gt;

&lt;p&gt;这也是为什么真正业务里的 Golden Set 往往比随便找一个 public benchmark 更有价值。我更愿意把 Golden Set 定义成：&lt;/p&gt;

&lt;blockquote&gt;
  &lt;p&gt;一组规模未必很大，但经过严格筛选、能代表真实 workload，并且拥有可信 grading criteria 的任务集合。&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;如果你做金融 Research Agent，那么 golden set 不应该是一堆「什么是 EBITDA」这样的金融知识题，而应该来自 analyst 真的会做的事情：从 earnings release 提取 guidance、重建 segment revenue、识别 GAAP/non-GAAP differences、从 10-K 中分析 debt maturity、根据 management commentary 判断 margin driver、检查 valuation model 中的数据引用。&lt;/p&gt;

&lt;p&gt;Public benchmark 回答的是 &lt;em&gt;How good is my model compared with everyone else?&lt;/em&gt;；Golden Set 回答的则是 &lt;em&gt;How good is my system at doing my job?&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;这两个问题根本不是一回事。&lt;/p&gt;

&lt;h2 id=&quot;04benchmark-最大的问题你最后会把考试本身学会&quot;&gt;04｜Benchmark 最大的问题：你最后会把考试本身学会&lt;/h2&gt;

&lt;p&gt;一个 benchmark 一旦被公开，就开始了一场不可避免的 race。研究者研究它，model developers 跑它，training data 可能包含它，post-training data 可能围绕它构造，prompt engineering 也会针对它优化。最终一个 benchmark 可能出现 saturation，甚至 contamination。&lt;/p&gt;

&lt;p&gt;这背后其实就是 Goodhart’s Law：&lt;/p&gt;

&lt;blockquote&gt;
  &lt;p&gt;&lt;strong&gt;When a measure becomes a target, it ceases to be a good measure.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;回到前面那个式子：我们优化的是可观测的 (M)，希望顺带提升不可观测的 (C)。但一旦优化压力足够大，模型完全可以只在 (\varepsilon) 上做文章——分数上去了，能力没动。&lt;/p&gt;

&lt;p&gt;如果所有团队都优化 MMLU，最终你很难判断模型到底变得更有 general intelligence，还是更擅长 MMLU-shaped tasks。&lt;/p&gt;

&lt;p&gt;更麻烦的是 contamination。公开题目长期存在于互联网后，就可能出现在 pretraining 或 post-training corpus 中。近年来甚至专门出现了 MMLU-CF 这样的工作，通过 closed test set 和 decontamination 规则试图减少 benchmark leakage；其出发点正是公开 MCQ benchmark 容易受到 data contamination 的影响。&lt;/p&gt;

&lt;p&gt;所以今天一个真正靠谱的 Evaluation System，通常不会押注在单一 static benchmark 上，而会同时使用 public benchmark、private benchmark、held-out golden set、dynamic task generation 和 production failure cases。&lt;/p&gt;

&lt;p&gt;&lt;mark&gt;Benchmark 不应该是一张毕业证，而应该是一支不断需要重新校准的温度计。&lt;/mark&gt;&lt;/p&gt;

&lt;h2 id=&quot;05开放式回答怎么办答案开始从对不对变成哪个好&quot;&gt;05｜开放式回答怎么办？答案开始从「对不对」变成「哪个好」&lt;/h2&gt;

&lt;p&gt;当 LLM 真正进入聊天、writing、analysis 之后，一个更加根本的问题出现了。&lt;/p&gt;

&lt;p&gt;假设用户说：「帮我写一封拒绝 offer 的邮件，但不要太冷漠。」模型 A 写得非常礼貌但啰嗦，模型 B 很简洁但稍微生硬。哪个「正确」？&lt;/p&gt;

&lt;p&gt;这里根本不存在 exact match。于是 Evaluation 从 answer correctness 进入 &lt;strong&gt;preference judgment&lt;/strong&gt;。&lt;/p&gt;

&lt;p&gt;Chatbot Arena 就是这个变化最典型的案例之一。它不规定一个所谓 golden answer，而是把两个匿名模型的输出放在一起，让真实用户做 pairwise comparison：A better、B better、tie。Chatbot Arena 的原始论文就是通过 crowdsourced pairwise human preferences 构建模型比较体系，并报告了用户投票与 expert raters 之间较好的 agreement。&lt;/p&gt;

&lt;p&gt;这个设计非常聪明，因为人其实不太擅长回答「这篇回答到底是 7.8 分还是 8.2 分」，但非常擅长回答「这两个里面哪个更好」。&lt;/p&gt;

&lt;p&gt;而 pairwise 之所以能变回一个可排序的分数，靠的是 Bradley–Terry 模型：给每个模型一个隐含实力 (s_i)，则&lt;/p&gt;

&lt;p&gt;[\Pr\big[\,i \succ j\,\big] \;=\; \frac{e^{s_i}}{e^{s_i} + e^{s_j}} \;=\; \sigma\big(s_i - s_j\big)]&lt;/p&gt;

&lt;p&gt;Elo 式的排行榜，本质上就是在用大量 pairwise 比较去反解这组 (s_i)。&lt;/p&gt;

&lt;p&gt;这也是为什么 pairwise preference 会同时出现在 Evaluation 和 Post-training 两个世界里。而这正是理解 RLHF 的入口。&lt;/p&gt;

&lt;h2 id=&quot;06rlhfevaluation-第一次直接变成-training-signal&quot;&gt;06｜RLHF：Evaluation 第一次直接变成 Training Signal&lt;/h2&gt;

&lt;p&gt;传统 supervised learning 的思路是：给模型一个 input，再给它一个 ideal output，让模型学习 imitation。但对于开放式 assistant，很多任务根本没有唯一的 ideal response。我们可能只知道：&lt;strong&gt;A 比 B 好&lt;/strong&gt;。&lt;/p&gt;

&lt;p&gt;RLHF——Reinforcement Learning from Human Feedback——做的核心事情，就是把这种 preference 转换成可以优化的 signal。&lt;/p&gt;

&lt;p&gt;InstructGPT 是最经典的例子之一。其 pipeline 大致可以理解成三步：先收集人类写出的 high-quality demonstrations 做 supervised fine-tuning；然后针对同一个 prompt 生成多个 candidate responses，让 human labelers 对回答进行 ranking；接着用这些 preference data 训练一个 reward model，再让语言模型通过 reinforcement learning 去最大化这个 reward。&lt;/p&gt;

&lt;p&gt;注意第二步用的正是上一节那个 Bradley–Terry 形式。reward model (r_\phi) 的训练目标是&lt;/p&gt;

&lt;p&gt;[\mathcal{L}(\phi) \;=\; -\,\mathbb{E}&lt;em&gt;{(x,\,y_w,\,y_l)\,\sim\,\mathcal{D}}\Big[\ \log \sigma\big(\, r&lt;/em&gt;\phi(x, y_w) - r_\phi(x, y_l) \,\big)\ \Big]]&lt;/p&gt;

&lt;p&gt;其中 (y_w) 是人类偏好的回答，(y_l) 是被拒绝的那个。然后 policy 的优化目标变成&lt;/p&gt;

&lt;p&gt;[\max_{\pi_\theta}\;\; \mathbb{E}&lt;em&gt;{x \sim \mathcal{D},\; y \sim \pi&lt;/em&gt;\theta(\cdot \mid x)}\big[\, r_\phi(x, y) \,\big] \;-\; \beta \, \mathbb{D}&lt;em&gt;{\mathrm{KL}}\Big[\, \pi&lt;/em&gt;\theta(y \mid x) \,\big|\, \pi_{\text{ref}}(y \mid x) \,\Big]]&lt;/p&gt;

&lt;p&gt;第二项那个 KL penalty 值得多看一眼。它的存在本身就是一句关于 evaluation 的坦白：&lt;strong&gt;我们并不完全相信这个 evaluator&lt;/strong&gt;。(\beta) 越小，policy 越敢于把 reward model 推到分布之外；而一旦推得太远，得到的往往不是更好的回答，而是 reward model 的漏洞。&lt;/p&gt;

&lt;p&gt;这里发生了一件非常关键的事情：&lt;/p&gt;

&lt;blockquote&gt;
  &lt;p&gt;&lt;strong&gt;Evaluator 不再只是测量模型，它开始塑造模型。&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Reward model 本质上就是一个 learned evaluator。于是一个非常自然的问题出现了：&lt;/p&gt;

&lt;blockquote&gt;
  &lt;p&gt;&lt;strong&gt;Who evaluates the evaluator?&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;如果 reward model 有 bias，policy 就会学习 exploit 这个 bias。如果 reward model 偏爱特别长的答案，模型就会越来越啰嗦；如果 RM 把 confident tone 错当成 correctness，模型就可能学会「更加自信地犯错」。&lt;/p&gt;

&lt;p&gt;所以 reward model 本身也需要 eval。这也是后来 RewardBench 这类 benchmark 出现的原因：reward model 已经成为 alignment pipeline 中的关键基础设施，因此必须单独测它判断 chosen / rejected response 的能力。&lt;/p&gt;

&lt;p&gt;Evaluation 从此开始递归：我们评模型；然后训练一个模型来评模型；然后还要再设计 benchmark 去评这个「评模型的模型」。&lt;/p&gt;

&lt;h2 id=&quot;07rlaif如果-human-feedback-太贵让-ai-自己当监督者呢&quot;&gt;07｜RLAIF：如果 Human Feedback 太贵，让 AI 自己当监督者呢？&lt;/h2&gt;

&lt;p&gt;RLHF 有一个非常现实的问题：human feedback 很贵，而且随着模型能力提升，人类会越来越难监督。如果模型正在证明一个复杂 theorem、分析几十万行代码、检查专业金融模型，一个普通 annotator 根本不知道答案到底好不好。&lt;/p&gt;

&lt;p&gt;于是自然出现了一个方向：&lt;strong&gt;Can AI supervise AI?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Anthropic 的 Constitutional AI 是这条路线的标志性工作之一。在其 RL 阶段，模型会生成 candidate responses，再由另一个模型根据预先定义的 constitutional principles 判断哪个回答更好，由这些 AI preferences 训练 preference model，再作为 reinforcement learning 的 reward signal；这就是 RLAIF——Reinforcement Learning from AI Feedback。&lt;/p&gt;

&lt;p&gt;这件事对 Evaluation 的意义远比「省人力」更大。因为一旦 evaluator 可以由 model scale，你就可以产生数量级更大的 feedback。但与此同时，新的风险也来了：如果 teacher model 和 student model 共享同样的 blind spot，整个 feedback loop 可能变成一种 self-reinforcing error。&lt;/p&gt;

&lt;p&gt;&lt;mark&gt;Human feedback 有 human bias，AI feedback 有 model bias。Evaluation 从来没有免费的午餐。&lt;/mark&gt;&lt;/p&gt;

&lt;h2 id=&quot;08llm-as-a-judgejudge-model-为什么好用又为什么危险&quot;&gt;08｜LLM-as-a-Judge：Judge Model 为什么好用，又为什么危险？&lt;/h2&gt;

&lt;p&gt;即使不做 RL，在日常产品开发里，团队也越来越频繁地使用 LLM-as-a-Judge。一个标准 judge prompt 通常包含 user question、reference context、candidate response 和 evaluation rubric，然后要求一个强模型输出 correctness、relevance、completeness、style 等 score。&lt;/p&gt;

&lt;p&gt;这非常 scalable，特别适合那些无法 deterministic grading 的任务。但 LLM judge 并不是 oracle。&lt;/p&gt;

&lt;p&gt;经典的 MT-Bench / LLM-as-a-Judge 研究发现，强模型作为 judge 可以与 human preference 达到较高 agreement，但同时也系统性存在 position bias、verbosity bias、self-enhancement bias 等问题。&lt;/p&gt;

&lt;p&gt;Position bias 很简单：把 A 放左边和把 A 放右边，judge 可能给出不同结果。检验它其实只要一个数：把顺序交换以后仍然给出同一结论的比例，&lt;/p&gt;

&lt;p&gt;[\text{consistency} \;=\; \Pr\Big[\, J(x,\,y_a,\,y_b) \;=\; \overline{J(x,\,y_b,\,y_a)} \,\Big]]&lt;/p&gt;

&lt;p&gt;如果这个数明显低于 1，那么你 leaderboard 上的差距里有一部分只是位置。Verbosity bias 则意味着一个很长、信息密度一般的回答，可能因为「看起来更完整」而赢过短而准确的答案。&lt;/p&gt;

&lt;p&gt;所以真正使用 LLM-as-a-Judge 时，不能只是「拿最强模型帮我打个分」。你至少要考虑：rubric 是否足够明确；用 absolute scoring 还是 pairwise；A/B 是否需要 swap position；judge 是否能看到 reference；是否要求 judge 给出 evidence；是否存在 domain blind spot；以及和 human expert 的 agreement 到底是多少。&lt;/p&gt;

&lt;p&gt;正确的思路不是把 LLM Judge 当成 ground truth，而是：&lt;/p&gt;

&lt;blockquote&gt;
  &lt;p&gt;&lt;strong&gt;Treat the judge as another noisy measurement instrument.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;先 calibration，再 scale。&lt;/p&gt;

&lt;h2 id=&quot;09outcome-reward-vs-process-reward只看最终答案够不够&quot;&gt;09｜Outcome Reward vs Process Reward：只看最终答案够不够？&lt;/h2&gt;

&lt;p&gt;Evaluation 继续向 reasoning 深处走，就会碰到另一个问题。&lt;/p&gt;

&lt;p&gt;假设一道数学题最终答案是 42。模型写了十步推理，第 4 步其实错了，第 7 步又阴差阳错地把错误抵消，最后结果刚好等于 42。按照 outcome-based evaluator：&lt;em&gt;Perfect.&lt;/em&gt; 按照我们真正希望模型具有的 reasoning：显然不应该 perfect。&lt;/p&gt;

&lt;p&gt;这就是 &lt;strong&gt;Outcome Reward Model（ORM）&lt;/strong&gt; 和 &lt;strong&gt;Process Reward Model（PRM）&lt;/strong&gt; 的区别。写出来，ORM 只看终点：&lt;/p&gt;

&lt;p&gt;[r_{\text{outcome}}\big(x,\, y_{1:T}\big) \;=\; \mathbf{1}\big{\, \text{answer}(y_{1:T}) = y^{\star} \,\big}]&lt;/p&gt;

&lt;p&gt;PRM 则对每一步 (y_t) 都给一个 step-level score (s_\phi(x, y_{1:t}))，再聚合，例如&lt;/p&gt;

&lt;p&gt;[r_{\text{process}}\big(x,\, y_{1:T}\big) \;=\; \min_{1 \le t \le T} s_\phi\big(x,\, y_{1:t}\big)
\qquad\text{或}\qquad
\prod_{t=1}^{T} s_\phi\big(x,\, y_{1:t}\big)]&lt;/p&gt;

&lt;p&gt;注意这里用 (\min) 或连乘而不是求平均，是有意为之：一条推理链的可信度应该由它最弱的一步决定，而不是被九个正确步骤平均掉。&lt;/p&gt;

&lt;p&gt;OpenAI 的 &lt;em&gt;Let’s Verify Step by Step&lt;/em&gt; 对这一问题做了系统实验。在其 MATH 实验中，process supervision 显著优于只根据最终结果进行监督的方法，并发布了包含约 80 万 step-level human feedback labels 的 PRM800K。&lt;/p&gt;

&lt;p&gt;为什么 PRM 这么重要？因为复杂 reasoning 里最大的问题并不是「最终有没有得到答案」，而是 error 可以在 trajectory 中不断累积。&lt;/p&gt;

&lt;p&gt;但 process supervision 也有自己的难题：一个问题可能存在很多条完全不同但同样合理的 reasoning path。如果你的 process rubric 过于 rigid，模型反而可能为了迎合 grader，失去探索 alternative reasoning strategy 的能力。&lt;/p&gt;

&lt;p&gt;因此到了 Agent 时代，一个越来越重要的原则会出现：&lt;/p&gt;

&lt;blockquote&gt;
  &lt;p&gt;&lt;strong&gt;Outcome first, process for diagnosis.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;最终有没有把事情办成，应该是主指标；trajectory evaluation 更多用于定位 failure，而不是强迫 agent 必须按照某一条「标准路径」行动。&lt;/p&gt;

&lt;h2 id=&quot;10best-of-nevaluator-还能在-inference-time-直接提高模型能力&quot;&gt;10｜Best-of-N：Evaluator 还能在 inference time 直接提高模型能力&lt;/h2&gt;

&lt;p&gt;Reward model 还有第三种用途，而且特别容易被忽略：它甚至不需要更新 model weights。&lt;/p&gt;

&lt;p&gt;假设模型面对一个问题一次性生成 (N) 个答案，再用 reward model 或 verifier 挑一个最好的：&lt;/p&gt;

&lt;p&gt;[y^{(1)}, \dots, y^{(N)} \;\overset{\text{i.i.d.}}{\sim}\; \pi_\theta(\,\cdot \mid x\,),
\qquad
\hat{y} \;=\; \arg\max_{1 \le i \le N}\; r_\phi\big(x,\, y^{(i)}\big)]&lt;/p&gt;

&lt;p&gt;这就是最简单的 &lt;strong&gt;Best-of-N（BoN）&lt;/strong&gt;。&lt;/p&gt;

&lt;p&gt;这里 evaluator 的角色又发生了变化：它既不是 benchmark，也不是 RL training signal，而是直接成为 inference-time search algorithm 的一部分。&lt;/p&gt;

&lt;p&gt;其实 HumanEval 当年 repeated sampling 能显著提高「至少找到一个正确程序」的比例，就已经展示了这种潜力。后来的 Best-of-N 方法则更加明确地使用 reward model 从多个 samples 中挑选最优候选。&lt;/p&gt;

&lt;p&gt;不过这里有一个必须写清楚的前提。设 (q(y)) 是我们真正关心的质量，(r_\phi) 只是它的 proxy。当 (N) 增大时，(\mathbb{E}\big[q(\hat{y})\big]) 是否随之上升，完全取决于 (r_\phi) 与 (q) 在&lt;strong&gt;分布尾部&lt;/strong&gt;是否仍然一致。搜索越激进，越容易挑中那些 (r_\phi) 很高、(q) 却并不高的样本——这就是 reward hacking 在 inference time 的版本。BoN 与 RLHF 里的 KL penalty，其实是在处理同一个问题的两种形式。&lt;/p&gt;

&lt;p&gt;这实际上揭示了今天所谓 test-time scaling 背后的一个核心规律：&lt;/p&gt;

&lt;blockquote&gt;
  &lt;p&gt;&lt;strong&gt;Generator 决定你能产生哪些 candidate；Evaluator 决定你能不能从里面找到好的那个。&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;模型越强，generator 当然重要；但当 sample 数量和 search depth 上升以后，verifier / reward model quality 会越来越接近整个系统的瓶颈。这就是为什么 Evaluation 和 Reasoning 的边界正在迅速消失。&lt;/p&gt;

&lt;h2 id=&quot;11agent-出现以后benchmark-从题目变成了环境&quot;&gt;11｜Agent 出现以后，Benchmark 从「题目」变成了「环境」&lt;/h2&gt;

&lt;p&gt;当 LLM 只是 chatbot 时，Evaluation 的基本结构非常简单：给一个输入 (x)，拿到一个输出 (y)，然后打分&lt;/p&gt;

&lt;p&gt;[\text{score} \;=\; g\big(y,\; y^{\star}\big)]&lt;/p&gt;

&lt;p&gt;但 Agent 完全不是这个结构。一个 Agent 的 trajectory 更像&lt;/p&gt;

&lt;p&gt;[\tau \;=\; \big(\, s_0,\; a_1,\; o_1,\; s_1,\; a_2,\; o_2,\; \dots,\; a_T,\; o_T,\; s_T \,\big)]&lt;/p&gt;

&lt;p&gt;其中 (a_t) 是 action（调用工具、点击页面、写文件），(o_t) 是 observation，(s_t) 是 environment state。而评分变成了对&lt;strong&gt;终态&lt;/strong&gt;的判断：&lt;/p&gt;

&lt;p&gt;[\text{success}(\tau) \;=\; \mathbf{1}\big{\, \Phi(s_T) \,\big}]&lt;/p&gt;

&lt;p&gt;也就是说，你检查的不再是模型写了什么，而是世界最后变成了什么样子。「最后回答写得好不好」甚至可能已经不是重点。&lt;/p&gt;

&lt;p&gt;比如一个 travel agent 的目标是「帮我找到符合条件的航班」。它可能要浏览网页、读取日期、过滤价格、比较行程、处理页面错误。你真正想评估的是：&lt;em&gt;Did the agent complete the task?&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;这就是 Agent Benchmark 出现的背景。AgentBench 直接把模型放进 8 个不同的 interactive environments，评估 reasoning 和 decision-making，而不仅仅是静态 QA。GAIA 则进一步把目标定义成 general AI assistant：其 466 个问题会要求 reasoning、multimodal understanding、web browsing 和 tool use，而且特意设计成「人觉得并不特别难，但 AI 系统很容易失败」的任务。原始论文中 human respondents 达到 92%，而当时配备 plugins 的 GPT-4 只有 15%，说明「考试题很强」与「真实 assistant 很稳健」完全是两个维度。&lt;/p&gt;

&lt;p&gt;WebArena 更进一步：它直接构造可以交互的真实感网站环境，包括 e-commerce、forum、software development、content management 等场景，然后检查 agent 是否真的完成了 web task。原论文的 best GPT-4-based agent end-to-end success rate 只有 14.41%，而 human performance 为 78.24%。&lt;/p&gt;

&lt;p&gt;注意这里发生的范式变化：&lt;/p&gt;

&lt;blockquote&gt;
  &lt;p&gt;传统 benchmark 给模型一张试卷；Agent benchmark 给模型一个世界。&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;而 evaluator 不再只是检查文本答案，还需要检查 environment state、tool invocation、task completion、side effects、constraint violations、trajectory、cost 和 latency。&lt;/p&gt;

&lt;p&gt;这也意味着未来最重要的 Evaluation infrastructure，很可能不是「题库」，而是 &lt;strong&gt;reproducible environment&lt;/strong&gt;。&lt;/p&gt;

&lt;h2 id=&quot;12agent-eval-真正难的地方failure-attribution&quot;&gt;12｜Agent Eval 真正难的地方：Failure Attribution&lt;/h2&gt;

&lt;p&gt;假设一个金融 Agent 最终把一家公司的 2026E EBITDA 算错了。一句「Answer incorrect」其实没有多大价值。真正有价值的是知道&lt;strong&gt;为什么&lt;/strong&gt;错。&lt;/p&gt;

&lt;p&gt;也许它搜索到了错误年份的 earnings report，这叫 retrieval failure；也许文档找对了但把 adjusted EBITDA 当成 GAAP operating income，这是 extraction / semantic failure；也许数字都正确但公式算错了，这是 calculation failure；也许分析完全正确，但 citation 引用了另一份文件，这是 citation failure；也许 external tool timeout 之后 agent 没有 retry，这是 recovery failure。&lt;/p&gt;

&lt;p&gt;因此成熟的 Agent Eval 最重要的东西之一不是总分，而是 &lt;strong&gt;Failure Taxonomy&lt;/strong&gt;：&lt;/p&gt;

&lt;table&gt;
  &lt;thead&gt;
    &lt;tr&gt;
      &lt;th&gt;Failure Type&lt;/th&gt;
      &lt;th&gt;Share&lt;/th&gt;
    &lt;/tr&gt;
  &lt;/thead&gt;
  &lt;tbody&gt;
    &lt;tr&gt;
      &lt;td&gt;Retrieval&lt;/td&gt;
      &lt;td&gt;27%&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Tool Use&lt;/td&gt;
      &lt;td&gt;19%&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Reasoning&lt;/td&gt;
      &lt;td&gt;18%&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Data Extraction&lt;/td&gt;
      &lt;td&gt;14%&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Calculation&lt;/td&gt;
      &lt;td&gt;9%&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Citation&lt;/td&gt;
      &lt;td&gt;8%&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Other&lt;/td&gt;
      &lt;td&gt;5%&lt;/td&gt;
    &lt;/tr&gt;
  &lt;/tbody&gt;
&lt;/table&gt;

&lt;p&gt;这张表的价值往往比「Overall score = 73.4」大得多，因为它直接回答了研发团队下一步应该改哪里。&lt;/p&gt;

&lt;p&gt;如果 40% 的 failure 都来自 retrieval，你继续做 reasoning RL 很可能没有太大意义；如果 agent 大多数时候都找到了正确资料，但计算环节不稳定，那么也许需要的是一个 deterministic calculator，而不是更大的模型。&lt;/p&gt;

&lt;p&gt;&lt;mark&gt;Eval 真正的作用并不是告诉你模型有多差，而是告诉你系统为什么差。&lt;/mark&gt;&lt;/p&gt;

&lt;h2 id=&quot;13为什么一个-overall-score-几乎永远不够&quot;&gt;13｜为什么一个 Overall Score 几乎永远不够？&lt;/h2&gt;

&lt;p&gt;假设一个新 checkpoint 的 overall score 从 82.1 升到 84.0。听起来很好。但如果拆开：&lt;/p&gt;

&lt;table&gt;
  &lt;thead&gt;
    &lt;tr&gt;
      &lt;th&gt;Slice&lt;/th&gt;
      &lt;th&gt;Old&lt;/th&gt;
      &lt;th&gt;New&lt;/th&gt;
    &lt;/tr&gt;
  &lt;/thead&gt;
  &lt;tbody&gt;
    &lt;tr&gt;
      &lt;td&gt;Math&lt;/td&gt;
      &lt;td&gt;81&lt;/td&gt;
      &lt;td&gt;89&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Coding&lt;/td&gt;
      &lt;td&gt;80&lt;/td&gt;
      &lt;td&gt;87&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Writing&lt;/td&gt;
      &lt;td&gt;84&lt;/td&gt;
      &lt;td&gt;85&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Finance&lt;/td&gt;
      &lt;td&gt;86&lt;/td&gt;
      &lt;td&gt;&lt;strong&gt;77&lt;/strong&gt;&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Safety&lt;/td&gt;
      &lt;td&gt;88&lt;/td&gt;
      &lt;td&gt;&lt;strong&gt;82&lt;/strong&gt;&lt;/td&gt;
    &lt;/tr&gt;
  &lt;/tbody&gt;
&lt;/table&gt;

&lt;p&gt;如果你做的是 Financial Copilot，这不是升级，是事故。&lt;/p&gt;

&lt;p&gt;所以真正的 Evaluation 必须做 &lt;strong&gt;slice analysis&lt;/strong&gt;。常见维度包括 task type、domain、difficulty、language、context length、tool type、risk level、user segment、failure class。&lt;/p&gt;

&lt;p&gt;与此同时还要看 statistical uncertainty。一个模型在 100 道题上 83%，另一个 84%，并不能自动推出后者更好——因为 eval 本身也是 sampling。二项比例的标准误是&lt;/p&gt;

&lt;p&gt;[\mathrm{SE} \;=\; \sqrt{\frac{\hat{p}\,(1 - \hat{p})}{n}}, \qquad
\text{95\% CI} \;=\; \hat{p} \;\pm\; 1.96\,\mathrm{SE}]&lt;/p&gt;

&lt;p&gt;代入 (\hat p = 0.83)、(n = 100)，(\mathrm{SE} \approx 3.8\%)，置信区间大约是 (\pm 7.4) 个百分点。也就是说，83% 和 84% 之间的差距，几乎完全淹没在噪声里。&lt;/p&gt;

&lt;p&gt;更好的做法是 paired comparison：在同一批题目上比较两个模型，只统计一个对、另一个错的那些题（McNemar 检验），这样可以消掉题目难度带来的方差，用同样的样本量得到高得多的分辨率。&lt;/p&gt;

&lt;p&gt;这就是为什么 leaderboard 上一个漂亮的 scalar score，在真实 model development 里往往只是入口。真正重要的是：&lt;/p&gt;

&lt;blockquote&gt;
  &lt;p&gt;&lt;strong&gt;Where did we improve, where did we regress, and why?&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2 id=&quot;14offline-evalonline-eval以及现实世界最后的一票&quot;&gt;14｜Offline Eval、Online Eval，以及现实世界最后的一票&lt;/h2&gt;

&lt;p&gt;再好的 golden set 都有一个天然问题：它只是现实世界的 proxy。最终产品成功与否，不是由 MMLU、SWE-bench 或某个 internal judge score 决定，而是由真实用户决定。&lt;/p&gt;

&lt;p&gt;于是 Eval 最后还要分成两个世界。&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Offline Eval&lt;/strong&gt; 用于开发阶段。它便宜、快速、可重复、可以 regression test，也能在上线前抓住明显问题。&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Online Eval&lt;/strong&gt; 则看 production 里的真实 outcome：task completion、user preference、regeneration rate、escalation rate、retention、conversion、latency、token cost、cost per successful task。&lt;/p&gt;

&lt;p&gt;最后一个指标值得单独写出来，因为它常常比单纯的 accuracy 更能反映系统的真实经济性：&lt;/p&gt;

&lt;p&gt;[\text{cost per successful task} \;=\; \frac{\mathbb{E}\big[\text{cost per attempt}\big]}{\Pr\big[\text{success}\big]}]&lt;/p&gt;

&lt;p&gt;一个把 success rate 从 60% 提到 75% 的改动，即使单次调用更贵，也可能整体更便宜。&lt;/p&gt;

&lt;p&gt;但这里又有一个坑：&lt;strong&gt;user preference 不等于 truth&lt;/strong&gt;。一个非常自信、语言漂亮、永远顺着用户说的模型，可能 immediate preference 很高，但 factuality 和 calibration 很差。因此真正的 product objective 往往是 multi-objective，写成带约束的形式会更诚实：&lt;/p&gt;

&lt;p&gt;[\max_{\text{system}} \;\; \mathbb{E}\big[\text{task success}\big]
\quad \text{s.t.} \quad
\text{factuality} \ge \tau_f,\;\;
\text{harm rate} \le \tau_s,\;\;
\mathbb{E}[\text{cost}] \le c,\;\;
p_{95}(\text{latency}) \le \ell]&lt;/p&gt;

&lt;p&gt;把 safety 和 factuality 放进约束而不是放进加权和，是有意义的：它们不应该被 success rate 的提升「买断」。&lt;/p&gt;

&lt;p&gt;这也是为什么根本不存在一个真正意义上的 &lt;em&gt;Universal LLM Score&lt;/em&gt;。&lt;mark&gt;模型能力不是 scalar，而是一个 vector。&lt;/mark&gt;所谓「哪个模型最好」，本身就是一个不完整的问题。真正的问题永远是：&lt;/p&gt;

&lt;blockquote&gt;
  &lt;p&gt;&lt;strong&gt;Best for what?&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2 id=&quot;15真正应该怎么搭一套-eval-system&quot;&gt;15｜真正应该怎么搭一套 Eval System？&lt;/h2&gt;

&lt;p&gt;如果今天从零开始做一个 LLM / Agent 产品，我不会先问「业界 benchmark 用什么」，而会按下面这套逻辑设计。&lt;/p&gt;

&lt;ol&gt;
  &lt;li&gt;&lt;strong&gt;收集真实任务&lt;/strong&gt;，而不是让 PM 在会议室里凭空造 prompts。你需要知道真正的 workload distribution 是什么。&lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;把任务做 taxonomy&lt;/strong&gt;：Task Type × Difficulty × Domain × Risk × Tool。&lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;抽出一个 high-quality Golden Set&lt;/strong&gt;。数量不一定特别大，但必须 representative，而且要包含 edge cases 和历史 production failures。&lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;为每一种任务设计 rubric&lt;/strong&gt;。什么叫 fully correct，什么叫 partial success，什么叫 critical failure，能不能 abstain，是否要求 citation，都要写清楚。&lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;能 deterministic grading 的绝不先上 LLM judge&lt;/strong&gt;。代码跑 tests，数字直接算，tool task 检查 final environment state。&lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;对无法直接执行的任务&lt;/strong&gt;（writing、analysis、open-ended QA），再引入 criteria-based LLM judge，并用 expert human labels 做 calibration。&lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;不只输出 overall score&lt;/strong&gt;，而是维护 slices 和 failure taxonomy。&lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;每一次更新都跑 regression suite&lt;/strong&gt;——model、prompt、retrieval、tool、system prompt，任何一处改动都算。&lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;让 production 中出现的新 failure 持续回流到 eval set。&lt;/strong&gt;&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;于是整个研发流程会变成一个闭环：&lt;/p&gt;

&lt;blockquote&gt;
  &lt;p&gt;real workload → golden set → rubric → grader → slice &amp;amp; failure analysis → 定位瓶颈 → 改 model / prompt / retrieval / tool → regression → 上线 → production failures → 回流到 golden set&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;这就是 &lt;strong&gt;Eval-driven Development&lt;/strong&gt;。&lt;/p&gt;

&lt;p&gt;它和传统 software engineering 里的 Test-driven Development 很像，但困难也大得多，因为传统软件测试的是 deterministic system，而 LLM 是 stochastic、开放式、甚至会主动与环境交互的系统。&lt;/p&gt;

&lt;h2 id=&quot;16eval-最后为什么会变成-ai-研发最核心的基础设施&quot;&gt;16｜Eval 最后为什么会变成 AI 研发最核心的基础设施？&lt;/h2&gt;

&lt;p&gt;现在我们可以重新看一遍整个历史。&lt;/p&gt;

&lt;p&gt;MMLU 时代，Evaluation 主要意味着&lt;strong&gt;给模型出题&lt;/strong&gt;。HumanEval 出现以后，我们发现&lt;strong&gt;不要只看文本，应该执行模型的产物&lt;/strong&gt;。SWE-bench 进一步告诉我们&lt;strong&gt;不要只测 isolated task，要测真实 workflow&lt;/strong&gt;。Chatbot Arena 告诉我们&lt;strong&gt;有些质量没有唯一答案，只能从 human preference 中学习&lt;/strong&gt;。&lt;/p&gt;

&lt;p&gt;RLHF 把 human preference 训练成 reward model，于是 &lt;strong&gt;Evaluation 开始成为 training objective&lt;/strong&gt;。RLAIF 让 AI 自己产生 preference，于是 &lt;strong&gt;evaluator 也开始 scale&lt;/strong&gt;。Process Reward 告诉我们，对于复杂 reasoning，可能不仅要评价 outcome，还要监督过程。Best-of-N 又把 evaluator 搬到了 inference time：模型生成很多可能性，verifier 决定哪个值得留下。最后 Agent benchmark 把整个问题推进到环境层面——我们不再评估「模型回答了什么」，而是在评估「系统究竟完成了什么」。&lt;/p&gt;

&lt;p&gt;所以今天再把 Evaluation 理解成「模型训练完以后跑几个 benchmark」，其实已经完全低估了它。在越来越多现代 AI system 中，Evaluator 同时承担至少四个角色：&lt;mark&gt;它测量模型，也塑造模型；它决定 reward，也指导 inference-time search。&lt;/mark&gt;它既告诉你哪个模型更强，也告诉你系统为什么失败。&lt;/p&gt;

&lt;p&gt;而当 foundation model 本身越来越容易获得——API 可以调用，open-weight model 可以下载，fine-tuning pipeline 越来越标准化——真正难复制的东西，反而可能变成 domain data、expert feedback、production failure history，以及长期积累下来的 evaluation infrastructure。&lt;/p&gt;

&lt;p&gt;因为 AI 开发里最危险的一件事情，从来不是「模型没有进步」，而是：&lt;/p&gt;

&lt;blockquote&gt;
  &lt;p&gt;&lt;strong&gt;你以为它进步了。&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;一个模型在 leaderboard 上涨了 5 分，却在你的核心用户任务上退化；一个 reward model score 一路上升，却只是越来越擅长 reward hacking；一个 Agent demo 看起来惊艳，但真正跑 1,000 个 production cases 时 success rate 一塌糊涂。&lt;/p&gt;

&lt;p&gt;没有好的 Eval，你甚至没有语言去描述这些问题。&lt;/p&gt;

&lt;p&gt;所以未来模型团队真正重要的问题，也许不会只是 &lt;em&gt;How do we train a smarter model?&lt;/em&gt;，而会越来越变成：&lt;/p&gt;

&lt;blockquote&gt;
  &lt;p&gt;&lt;strong&gt;What does “smarter” actually mean, and how do we know when we get there?&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;这就是 Evaluation。它表面上是在给 AI 打分；实际上，它是在定义我们究竟想把 AI 变成什么。&lt;/p&gt;
</content>
  </entry>
  
  <entry>
    <title>How Do We Actually Know a Model Is Getting Better?</title>
    <link href="https://mochiaochen.github.io/en/writing/2026/08/how-we-know-a-model-is-better/" rel="alternate" type="text/html"/>
    <published>2026-08-29T10:00:00+08:00</published>
    <updated>2026-08-29T10:00:00+08:00</updated>
    <id>https://mochiaochen.github.io/en/writing/2026/08/how-we-know-a-model-is-better-en</id>
    <content type="html" xml:base="https://mochiaochen.github.io/en/writing/2026/08/how-we-know-a-model-is-better/">&lt;p&gt;Compress the last few years of frontier model competition into a sequence of questions and you get something like this. First everyone asked &lt;em&gt;How big is the model?&lt;/em&gt; Then it became &lt;em&gt;How much compute did you use?&lt;/em&gt; Then &lt;em&gt;What is your MMLU score?&lt;/em&gt; And now that models are genuinely entering coding, research, finance, customer support—and driving computers and browsers directly—a more awkward question has surfaced:&lt;/p&gt;

&lt;blockquote&gt;
  &lt;p&gt;&lt;strong&gt;How do we actually know whether the model is getting better?&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;This is evaluation, or, as most people say, eval.&lt;/p&gt;

&lt;p&gt;On the surface eval is easy to describe: give the model a set of questions and compute a score. But anyone who has actually worked on model development discovers quickly that evaluation may be the most underrated part of the whole LLM pipeline, and also the part closest to the centre of it. It is not merely a QA system you run after training to sign the model off. &lt;mark&gt;How you define “good” ends up determining what data the team collects, how post-training is done, what the reward model learns, what RL optimises, and even what the system should search over at inference time.&lt;/mark&gt;&lt;/p&gt;

&lt;p&gt;Put differently:&lt;/p&gt;

&lt;blockquote&gt;
  &lt;p&gt;&lt;strong&gt;The model you get is downstream of the eval you build.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;In that sense the real starting point of model development is not training. It is measurement.&lt;/p&gt;

&lt;h2 id=&quot;01eval-is-not-a-benchmark-first-work-out-what-you-are-measuring&quot;&gt;01｜Eval is not a benchmark: first work out what you are measuring&lt;/h2&gt;

&lt;p&gt;Most people first meet evaluation as a jumble of words—benchmark, test set, leaderboard, eval. In fact a benchmark is only one component. A complete evaluation system can be written, at minimum, as a six-tuple:&lt;/p&gt;

&lt;p&gt;[\text{Eval} \;=\; \big\langle\, C,\; T,\; G,\; J,\; M,\; P \,\big\rangle]&lt;/p&gt;

&lt;p&gt;Here (C) is the construct you want to measure (the capability itself), (T) is the task distribution meant to represent it, (G) is the ground truth or rubric, (J) is the judge, (M) is the metric that turns judgements into numbers, and (P) is the protocol under which the model is tested—prompt template, number of few-shot examples, temperature, tool access, context length.&lt;/p&gt;

&lt;p&gt;In other words, you owe six answers: which capability? which tasks stand in for it? what counts as a good response? who decides? how does a decision become a number? and under what conditions is the model tested?&lt;/p&gt;

&lt;p&gt;Take MMLU, the canonical case. It puts the model into a highly standardised examination setting: 57 tasks spanning elementary mathematics, US history, computer science, law and more, scored by multiple-choice accuracy across broad knowledge and problem solving. When it was proposed in 2020, models were still far from expert-level performance on these tasks, which is exactly what made it good at separating them.&lt;/p&gt;

&lt;p&gt;And here the single most important concept in evaluation appears: &lt;strong&gt;construct validity&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;What does a high MMLU score actually tell you? At minimum, that the model does well on this particular set of multi-subject multiple-choice questions. Does it mean the model is “smarter”? That it makes a better research assistant? That it codes better? That it plans better in the real world? None of that follows.&lt;/p&gt;

&lt;p&gt;Which gives the first principle of all evaluation:&lt;/p&gt;

&lt;blockquote&gt;
  &lt;p&gt;&lt;strong&gt;Never confuse the metric with the construct.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;A metric is a proxy designed to measure some abstract capability. It is not the capability. Written out: what we care about is (C), but all we can observe is (M), and there is always a layer in between:&lt;/p&gt;

&lt;p&gt;[M \;=\; f(C) + \varepsilon]&lt;/p&gt;

&lt;p&gt;where (\varepsilon) absorbs task-sampling bias, grader noise, sensitivity to prompt formatting, and data contamination. Most of the work in eval engineering is really about shrinking (\varepsilon)—and being honest that it never reaches zero.&lt;/p&gt;

&lt;p&gt;People say they want to measure reasoning. But reasoning decomposes into at least deductive reasoning, mathematical reasoning, causal reasoning, multi-hop reasoning, planning, counterfactual reasoning, long-horizon reasoning. If you have not defined which one you mean, then no amount of dataset size, grader sophistication, or statistical elegance downstream will save you.&lt;/p&gt;

&lt;p&gt;A mature eval therefore never begins with “which benchmark should we run.” It begins with a sentence:&lt;/p&gt;

&lt;blockquote&gt;
  &lt;p&gt;&lt;strong&gt;What capability or behaviour are we trying to measure?&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2 id=&quot;02from-mmlu-to-humaneval-why-a-correct-answer-is-not-enough&quot;&gt;02｜From MMLU to HumanEval: why a “correct answer” is not enough&lt;/h2&gt;

&lt;p&gt;MMLU represents a classical family of benchmarks: static questions, determinate answers, and a very simple scoring function—exact match or accuracy.&lt;/p&gt;

&lt;p&gt;This kind of eval has an enormous engineering advantage: it is clean. Hold the protocol fixed, and comparing model A at 80% with model B at 85% is relatively easy.&lt;/p&gt;

&lt;p&gt;But once a model moves from answering questions to producing something that can be executed, string matching starts to fail.&lt;/p&gt;

&lt;p&gt;In 2021 OpenAI introduced HumanEval in the Codex paper. The key change was not that the questions became programming problems; it was that the scoring philosophy changed. The model generates a Python function from a docstring, and what is evaluated is &lt;strong&gt;functional correctness&lt;/strong&gt;—whether the code actually works, not whether the text resembles some reference answer. In the original paper Codex solved 28.8% of HumanEval problems, and, in a result that turned out to matter a great deal, allowing 100 samples per problem raised the share for which at least one working solution was found to 70.2%.&lt;/p&gt;

&lt;p&gt;That single result anticipated two lines of work that have only grown more important.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;First: for executable tasks, the environment is the best evaluator you have.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;If you ask “is this code any good,” another LLM can certainly give an opinion. But if the real question is “does this code satisfy the requirement,” the most trustworthy method is usually still: &lt;em&gt;run the tests.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;This is deterministic eval—exact match, unit tests, compiler and runtime, SQL execution, schema validation, numerical tolerance, state checking.&lt;/p&gt;

&lt;p&gt;&lt;mark&gt;Where a task has objective, executable ground truth, a deterministic evaluator is usually more reliable than an LLM judge.&lt;/mark&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Second, and more interesting: HumanEval showed that model capability is not a single deterministic output but a probability distribution.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;For a given problem (x), what we actually face is&lt;/p&gt;

&lt;p&gt;[y \;\sim\; p_\theta(\,\cdot \mid x\,)]&lt;/p&gt;

&lt;p&gt;so the probability of being right on the first attempt is&lt;/p&gt;

&lt;p&gt;[\text{pass@}1 \;=\; \mathbb{E}&lt;em&gt;{x}\Big[\ \mathbb{E}&lt;/em&gt;{y \sim p_\theta(\cdot\mid x)}\big[\ \mathbf{1}{\text{correct}(y)}\ \big]\Big]]&lt;/p&gt;

&lt;p&gt;and “can the model find a correct answer given (k) attempts” is a different quantity entirely. Sampling (n) candidates per problem of which (c) are correct, the Codex paper’s unbiased estimator is&lt;/p&gt;

&lt;p&gt;[\text{pass@}k \;=\; \mathbb{E}_{x}\left[\, 1 - \frac{\dbinom{n-c}{k}}{\dbinom{n}{k}} \,\right]]&lt;/p&gt;

&lt;p&gt;This line runs all the way to Best-of-N, self-consistency, search, verifiers, and test-time compute.&lt;/p&gt;

&lt;p&gt;So from HumanEval onwards we stopped only asking:&lt;/p&gt;

&lt;blockquote&gt;
  &lt;p&gt;&lt;strong&gt;Can the model answer this question?&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;and started asking:&lt;/p&gt;

&lt;blockquote&gt;
  &lt;p&gt;&lt;strong&gt;Can the system search its output space until it finds a correct answer?&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2 id=&quot;03from-humaneval-to-swe-bench-writing-code-is-not-doing-software-engineering&quot;&gt;03｜From HumanEval to SWE-bench: writing code is not doing software engineering&lt;/h2&gt;

&lt;p&gt;HumanEval settled something important: do not compare code to a reference for resemblance, test it for functional correctness. But it retains an obvious limitation—it mostly tests relatively self-contained function generation.&lt;/p&gt;

&lt;p&gt;Real software engineering does not look like that. What a programmer actually receives is closer to: “there’s a GitHub issue on this repository, a user says pagination breaks under some condition, go fix it.”&lt;/p&gt;

&lt;p&gt;The model must understand the issue, read a repository that may run to tens of thousands of lines, locate the relevant files, understand the existing architecture, modify one or more files, run the tests, and avoid introducing a regression.&lt;/p&gt;

&lt;p&gt;That is what SWE-bench set out to measure. The original benchmark collected 2,294 real GitHub issues and their corresponding pull requests from 12 real Python repositories. The model is handed the codebase and the issue description, and the task is not to write a function but to modify the repository so the issue is genuinely resolved—often requiring coordinated edits across functions, classes, and files. When the paper first appeared, even the best systems solved only a small fraction.&lt;/p&gt;

&lt;p&gt;This is a large migration in benchmark design philosophy:&lt;/p&gt;

&lt;blockquote&gt;
  &lt;p&gt;HumanEval measures &lt;strong&gt;code generation&lt;/strong&gt;. SWE-bench measures &lt;strong&gt;software engineering task completion&lt;/strong&gt;.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;SWE-bench Verified went a step further and added human validation: 500 instances screened by people to ensure the issue description is clear, the test patch is reasonable, and the task is genuinely solvable from the information given.&lt;/p&gt;

&lt;p&gt;Which surfaces a problem in eval engineering that is routinely ignored: &lt;strong&gt;benchmarks have bugs too&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;If a task is impossible, or the reference answer is wrong, or the grader is miswritten, then what you measure is not model capability but dataset noise. High-quality eval is therefore not a matter of collecting more questions; it requires continuous dataset validation, error analysis, human audit, and versioning.&lt;/p&gt;

&lt;p&gt;This is also why, inside a real business, a golden set is usually worth more than any public benchmark you might pick up. I would define a golden set as:&lt;/p&gt;

&lt;blockquote&gt;
  &lt;p&gt;a set of tasks, not necessarily large, that has been rigorously screened, represents the real workload, and comes with trustworthy grading criteria.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;If you are building a financial research agent, your golden set should not be a pile of “what is EBITDA” knowledge questions. It should come from what analysts actually do: extract guidance from an earnings release, rebuild segment revenue, identify GAAP versus non-GAAP differences, analyse debt maturity from a 10-K, infer margin drivers from management commentary, check the data references inside a valuation model.&lt;/p&gt;

&lt;p&gt;A public benchmark answers &lt;em&gt;How good is my model compared with everyone else?&lt;/em&gt; A golden set answers &lt;em&gt;How good is my system at doing my job?&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;These are not the same question.&lt;/p&gt;

&lt;h2 id=&quot;04the-deepest-problem-with-benchmarks-eventually-you-learn-the-exam&quot;&gt;04｜The deepest problem with benchmarks: eventually you learn the exam&lt;/h2&gt;

&lt;p&gt;The moment a benchmark is published, an unavoidable race begins. Researchers study it, model developers run it, training data may contain it, post-training data may be constructed around it, prompt engineering is tuned against it. Eventually the benchmark saturates, or becomes contaminated.&lt;/p&gt;

&lt;p&gt;Underneath this is Goodhart’s Law:&lt;/p&gt;

&lt;blockquote&gt;
  &lt;p&gt;&lt;strong&gt;When a measure becomes a target, it ceases to be a good measure.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Return to the earlier expression. We optimise the observable (M) in the hope of raising the unobservable (C). But once optimisation pressure is high enough, a model can simply work on (\varepsilon) instead—the score rises and the capability does not move.&lt;/p&gt;

&lt;p&gt;If every team optimises MMLU, it becomes very hard to tell whether models are acquiring more general intelligence or merely getting better at MMLU-shaped tasks.&lt;/p&gt;

&lt;p&gt;Contamination is worse. Once public questions have lived on the internet for long enough, they may appear in a pretraining or post-training corpus. Work such as MMLU-CF has appeared specifically to reduce benchmark leakage through closed test sets and decontamination rules, precisely because public multiple-choice benchmarks are so exposed to it.&lt;/p&gt;

&lt;p&gt;A serious evaluation system today therefore does not bet on a single static benchmark. It uses public benchmarks, private benchmarks, held-out golden sets, dynamic task generation, and production failure cases together.&lt;/p&gt;

&lt;p&gt;&lt;mark&gt;A benchmark should not be a diploma. It should be a thermometer that needs recalibrating.&lt;/mark&gt;&lt;/p&gt;

&lt;h2 id=&quot;05what-about-open-ended-answers-from-is-it-right-to-which-is-better&quot;&gt;05｜What about open-ended answers? From “is it right” to “which is better”&lt;/h2&gt;

&lt;p&gt;Once LLMs entered chat, writing, and analysis, a more fundamental problem appeared.&lt;/p&gt;

&lt;p&gt;Suppose the user says: “write me an email declining an offer, but don’t make it cold.” Model A is very polite but long-winded. Model B is concise but slightly blunt. Which is &lt;em&gt;correct&lt;/em&gt;?&lt;/p&gt;

&lt;p&gt;There is no exact match here. So evaluation moves from answer correctness to &lt;strong&gt;preference judgement&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Chatbot Arena is the clearest example of the shift. It prescribes no golden answer; it places the outputs of two anonymous models side by side and asks real users for a pairwise comparison: A better, B better, tie. The original paper builds a model-comparison system out of crowdsourced pairwise human preferences and reports reasonable agreement between crowd votes and expert raters.&lt;/p&gt;

&lt;p&gt;The design is clever because people are not good at answering “is this response a 7.8 or an 8.2,” but are very good at answering “which of these two is better.”&lt;/p&gt;

&lt;p&gt;The bridge back from pairwise comparisons to an orderable score is the Bradley–Terry model: give each model a latent strength (s_i), and&lt;/p&gt;

&lt;p&gt;[\Pr\big[\,i \succ j\,\big] \;=\; \frac{e^{s_i}}{e^{s_i} + e^{s_j}} \;=\; \sigma\big(s_i - s_j\big)]&lt;/p&gt;

&lt;p&gt;An Elo-style leaderboard is, in essence, solving for these (s_i) from a large number of pairwise comparisons.&lt;/p&gt;

&lt;p&gt;This is also why pairwise preference shows up in evaluation and in post-training at the same time—and it is the entry point to understanding RLHF.&lt;/p&gt;

&lt;h2 id=&quot;06rlhf-the-first-time-evaluation-becomes-a-training-signal-directly&quot;&gt;06｜RLHF: the first time evaluation becomes a training signal directly&lt;/h2&gt;

&lt;p&gt;Supervised learning says: give the model an input, give it an ideal output, and have it imitate. But for an open-ended assistant many tasks have no unique ideal response. Often all we know is: &lt;strong&gt;A is better than B&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;RLHF—reinforcement learning from human feedback—is the machinery that turns that preference into something optimisable.&lt;/p&gt;

&lt;p&gt;InstructGPT is the canonical example. The pipeline is roughly three steps: collect high-quality human demonstrations for supervised fine-tuning; generate multiple candidate responses for the same prompt and have human labellers rank them; train a reward model on that preference data, and then use reinforcement learning to have the language model maximise that reward.&lt;/p&gt;

&lt;p&gt;Note that the second step uses exactly the Bradley–Terry form from the previous section. The training objective for the reward model (r_\phi) is&lt;/p&gt;

&lt;p&gt;[\mathcal{L}(\phi) \;=\; -\,\mathbb{E}&lt;em&gt;{(x,\,y_w,\,y_l)\,\sim\,\mathcal{D}}\Big[\ \log \sigma\big(\, r&lt;/em&gt;\phi(x, y_w) - r_\phi(x, y_l) \,\big)\ \Big]]&lt;/p&gt;

&lt;p&gt;where (y_w) is the preferred response and (y_l) the rejected one. The policy objective then becomes&lt;/p&gt;

&lt;p&gt;[\max_{\pi_\theta}\;\; \mathbb{E}&lt;em&gt;{x \sim \mathcal{D},\; y \sim \pi&lt;/em&gt;\theta(\cdot \mid x)}\big[\, r_\phi(x, y) \,\big] \;-\; \beta \, \mathbb{D}&lt;em&gt;{\mathrm{KL}}\Big[\, \pi&lt;/em&gt;\theta(y \mid x) \,\big|\, \pi_{\text{ref}}(y \mid x) \,\Big]]&lt;/p&gt;

&lt;p&gt;That KL penalty deserves a second look. Its very presence is a confession about evaluation: &lt;strong&gt;we do not fully trust this evaluator&lt;/strong&gt;. The smaller (\beta) is, the further the policy dares to push the reward model out of distribution—and past a certain point what you get is not better responses but the reward model’s blind spots.&lt;/p&gt;

&lt;p&gt;Something important has happened here:&lt;/p&gt;

&lt;blockquote&gt;
  &lt;p&gt;&lt;strong&gt;The evaluator no longer merely measures the model. It has begun to shape it.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;A reward model is, at bottom, a learned evaluator. Which raises the obvious question:&lt;/p&gt;

&lt;blockquote&gt;
  &lt;p&gt;&lt;strong&gt;Who evaluates the evaluator?&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;If the reward model has a bias, the policy will learn to exploit it. If the reward model prefers long answers, the model becomes more verbose. If it mistakes a confident tone for correctness, the model learns to be wrong more confidently.&lt;/p&gt;

&lt;p&gt;So reward models need eval of their own. This is why benchmarks such as RewardBench appeared: the reward model has become critical infrastructure in the alignment pipeline, and its ability to separate chosen from rejected responses has to be measured separately.&lt;/p&gt;

&lt;p&gt;Evaluation has become recursive. We evaluate models; then we train a model to evaluate models; then we design a benchmark to evaluate the model that evaluates models.&lt;/p&gt;

&lt;h2 id=&quot;07rlaif-if-human-feedback-is-expensive-can-ai-supervise-itself&quot;&gt;07｜RLAIF: if human feedback is expensive, can AI supervise itself?&lt;/h2&gt;

&lt;p&gt;RLHF has a practical problem: human feedback is expensive, and as models improve, humans get worse at supervising them. If a model is proving a difficult theorem, analysing hundreds of thousands of lines of code, or checking a professional financial model, an ordinary annotator simply cannot tell whether the answer is good.&lt;/p&gt;

&lt;p&gt;Hence the natural question: &lt;strong&gt;can AI supervise AI?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Anthropic’s Constitutional AI is a landmark on this path. In its RL stage, the model generates candidate responses and another model judges which is better against a set of pre-defined constitutional principles; those AI preferences train a preference model that becomes the reward signal for reinforcement learning. This is RLAIF—reinforcement learning from AI feedback.&lt;/p&gt;

&lt;p&gt;The significance for evaluation goes well beyond saving labour. Once the evaluator itself can be scaled by a model, you can produce orders of magnitude more feedback. But a new risk arrives with it: if the teacher model and the student model share the same blind spot, the entire feedback loop can become a self-reinforcing error.&lt;/p&gt;

&lt;p&gt;&lt;mark&gt;Human feedback carries human bias; AI feedback carries model bias. There is no free lunch in evaluation.&lt;/mark&gt;&lt;/p&gt;

&lt;h2 id=&quot;08llm-as-a-judge-why-judge-models-are-so-useful-and-so-dangerous&quot;&gt;08｜LLM-as-a-judge: why judge models are so useful, and so dangerous&lt;/h2&gt;

&lt;p&gt;Even without RL, teams increasingly reach for LLM-as-a-judge in everyday product work. A standard judge prompt contains a user question, reference context, a candidate response, and an evaluation rubric, and asks a strong model to output scores for correctness, relevance, completeness, style.&lt;/p&gt;

&lt;p&gt;It is extremely scalable, and it suits tasks that cannot be graded deterministically. But an LLM judge is not an oracle.&lt;/p&gt;

&lt;p&gt;The MT-Bench / LLM-as-a-judge work found that strong models as judges can reach high agreement with human preference, while also exhibiting systematic position bias, verbosity bias, and self-enhancement bias.&lt;/p&gt;

&lt;p&gt;Position bias is simple: put A on the left and put A on the right, and the judge may decide differently. Testing for it takes a single number—the fraction of pairs on which the judge reaches the same conclusion after the order is swapped:&lt;/p&gt;

&lt;p&gt;[\text{consistency} \;=\; \Pr\Big[\, J(x,\,y_a,\,y_b) \;=\; \overline{J(x,\,y_b,\,y_a)} \,\Big]]&lt;/p&gt;

&lt;p&gt;If that number is meaningfully below 1, some of the gap on your leaderboard is just position. Verbosity bias means a long answer of ordinary information density can beat a short, accurate one because it &lt;em&gt;looks&lt;/em&gt; more complete.&lt;/p&gt;

&lt;p&gt;So using an LLM judge properly is not “ask the strongest model for a score.” At minimum you have to consider: whether the rubric is specific enough; absolute scoring or pairwise; whether A/B positions need swapping; whether the judge can see a reference; whether the judge must cite evidence; whether it has a domain blind spot; and what its agreement with human experts actually is.&lt;/p&gt;

&lt;p&gt;The right posture is not to treat the judge as ground truth but:&lt;/p&gt;

&lt;blockquote&gt;
  &lt;p&gt;&lt;strong&gt;Treat the judge as another noisy measurement instrument.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Calibrate first, then scale.&lt;/p&gt;

&lt;h2 id=&quot;09outcome-reward-vs-process-reward-is-the-final-answer-enough&quot;&gt;09｜Outcome reward vs process reward: is the final answer enough?&lt;/h2&gt;

&lt;p&gt;Push evaluation deeper into reasoning and another problem appears.&lt;/p&gt;

&lt;p&gt;Suppose the answer to a maths problem is 42. The model writes ten steps of reasoning; step 4 is actually wrong; step 7 happens to cancel the error; the final result comes out at exactly 42. To an outcome-based evaluator: &lt;em&gt;perfect&lt;/em&gt;. To the kind of reasoning we actually want: clearly not.&lt;/p&gt;

&lt;p&gt;This is the distinction between an &lt;strong&gt;outcome reward model (ORM)&lt;/strong&gt; and a &lt;strong&gt;process reward model (PRM)&lt;/strong&gt;. Written out, the ORM looks only at the endpoint:&lt;/p&gt;

&lt;p&gt;[r_{\text{outcome}}\big(x,\, y_{1:T}\big) \;=\; \mathbf{1}\big{\, \text{answer}(y_{1:T}) = y^{\star} \,\big}]&lt;/p&gt;

&lt;p&gt;A PRM assigns a step-level score (s_\phi(x, y_{1:t})) to each step and aggregates, for instance as&lt;/p&gt;

&lt;p&gt;[r_{\text{process}}\big(x,\, y_{1:T}\big) \;=\; \min_{1 \le t \le T} s_\phi\big(x,\, y_{1:t}\big)
\qquad\text{or}\qquad
\prod_{t=1}^{T} s_\phi\big(x,\, y_{1:t}\big)]&lt;/p&gt;

&lt;p&gt;The choice of a minimum or a product rather than a mean is deliberate: a chain of reasoning should be only as trustworthy as its weakest step, not have that step averaged away by nine sound ones.&lt;/p&gt;

&lt;p&gt;OpenAI’s &lt;em&gt;Let’s Verify Step by Step&lt;/em&gt; studied this systematically. In its MATH experiments, process supervision substantially outperformed supervision based on the final result alone, and the work released PRM800K, roughly 800,000 step-level human feedback labels.&lt;/p&gt;

&lt;p&gt;Why does this matter? Because the hard part of complex reasoning is not whether an answer eventually appears, but that errors accumulate along the trajectory.&lt;/p&gt;

&lt;p&gt;Process supervision has its own difficulty, though: a problem may admit many different but equally valid reasoning paths. If your process rubric is too rigid, the model may sacrifice the exploration of alternative strategies in order to please the grader.&lt;/p&gt;

&lt;p&gt;Hence a principle that matters more and more in the agent era:&lt;/p&gt;

&lt;blockquote&gt;
  &lt;p&gt;&lt;strong&gt;Outcome first, process for diagnosis.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Whether the job actually got done should be the headline metric; trajectory evaluation is mostly for locating failures, not for forcing an agent down one canonical path.&lt;/p&gt;

&lt;h2 id=&quot;10best-of-n-the-evaluator-can-raise-capability-at-inference-time&quot;&gt;10｜Best-of-N: the evaluator can raise capability at inference time&lt;/h2&gt;

&lt;p&gt;Reward models have a third use that is easy to overlook: it does not require updating any weights.&lt;/p&gt;

&lt;p&gt;Suppose the model produces (N) answers to a question in one go, and a reward model or verifier picks the best:&lt;/p&gt;

&lt;p&gt;[y^{(1)}, \dots, y^{(N)} \;\overset{\text{i.i.d.}}{\sim}\; \pi_\theta(\,\cdot \mid x\,),
\qquad
\hat{y} \;=\; \arg\max_{1 \le i \le N}\; r_\phi\big(x,\, y^{(i)}\big)]&lt;/p&gt;

&lt;p&gt;This is Best-of-N in its simplest form.&lt;/p&gt;

&lt;p&gt;The evaluator’s role has changed again. It is neither a benchmark nor an RL training signal; it has become part of an inference-time search algorithm.&lt;/p&gt;

&lt;p&gt;HumanEval’s repeated sampling result already hinted at the potential. Later Best-of-N methods made it explicit by using a reward model to select among many samples.&lt;/p&gt;

&lt;p&gt;There is a precondition worth stating plainly. Let (q(y)) be the quality we actually care about and (r_\phi) a proxy for it. Whether (\mathbb{E}\big[q(\hat{y})\big]) rises with (N) depends entirely on whether (r_\phi) and (q) still agree &lt;strong&gt;in the tail&lt;/strong&gt;. The more aggressive the search, the easier it is to select samples where (r_\phi) is high and (q) is not—reward hacking, in its inference-time form. Best-of-N and the KL penalty in RLHF are two treatments of the same disease.&lt;/p&gt;

&lt;p&gt;Which reveals the rule underneath what is now called test-time scaling:&lt;/p&gt;

&lt;blockquote&gt;
  &lt;p&gt;&lt;strong&gt;The generator determines which candidates can exist. The evaluator determines whether you can find the good one among them.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The stronger the model, the more the generator matters—but as sample counts and search depth rise, verifier and reward model quality move steadily closer to being the bottleneck of the whole system. This is why the boundary between evaluation and reasoning is disappearing.&lt;/p&gt;

&lt;h2 id=&quot;11with-agents-the-benchmark-stops-being-an-exam-and-becomes-an-environment&quot;&gt;11｜With agents, the benchmark stops being an exam and becomes an environment&lt;/h2&gt;

&lt;p&gt;When an LLM is just a chatbot, the structure of evaluation is simple: take an input (x), get an output (y), and score it&lt;/p&gt;

&lt;p&gt;[\text{score} \;=\; g\big(y,\; y^{\star}\big)]&lt;/p&gt;

&lt;p&gt;An agent is nothing like this. Its trajectory looks more like&lt;/p&gt;

&lt;p&gt;[\tau \;=\; \big(\, s_0,\; a_1,\; o_1,\; s_1,\; a_2,\; o_2,\; \dots,\; a_T,\; o_T,\; s_T \,\big)]&lt;/p&gt;

&lt;p&gt;where (a_t) is an action (a tool call, a click, a file write), (o_t) an observation, and (s_t) the environment state. Scoring becomes a judgement about the &lt;strong&gt;final state&lt;/strong&gt;:&lt;/p&gt;

&lt;p&gt;[\text{success}(\tau) \;=\; \mathbf{1}\big{\, \Phi(s_T) \,\big}]&lt;/p&gt;

&lt;p&gt;You are no longer checking what the model wrote. You are checking what the world looks like afterwards. How well the closing message is phrased may not even be the point.&lt;/p&gt;

&lt;p&gt;A travel agent’s goal is “find me a flight that meets these conditions.” It may need to browse pages, read dates, filter prices, compare itineraries, handle page errors. What you want to evaluate is: &lt;em&gt;did the agent complete the task?&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;This is the setting in which agent benchmarks appeared. AgentBench places models into 8 distinct interactive environments to evaluate reasoning and decision-making rather than static QA. GAIA goes further and defines the target as a general AI assistant: its 466 questions require reasoning, multimodal understanding, web browsing, and tool use, and are deliberately designed to be tasks humans do not find especially hard but AI systems fail easily. In the original paper human respondents reached 92% while GPT-4 with plugins reached 15%—evidence that “strong on exams” and “robust as an assistant” are entirely different dimensions.&lt;/p&gt;

&lt;p&gt;WebArena goes further still, constructing interactive, realistic websites—e-commerce, forum, software development, content management—and checking whether the agent truly completed a web task. In the original paper the best GPT-4-based agent achieved an end-to-end success rate of 14.41% against human performance of 78.24%.&lt;/p&gt;

&lt;p&gt;Note the paradigm shift:&lt;/p&gt;

&lt;blockquote&gt;
  &lt;p&gt;A traditional benchmark hands the model an exam paper. An agent benchmark hands it a world.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;And the evaluator no longer checks a text answer alone. It has to check environment state, tool invocation, task completion, side effects, constraint violations, trajectory, cost, and latency.&lt;/p&gt;

&lt;p&gt;Which suggests that the most important evaluation infrastructure of the next few years is not a question bank but a &lt;strong&gt;reproducible environment&lt;/strong&gt;.&lt;/p&gt;

&lt;h2 id=&quot;12the-genuinely-hard-part-of-agent-eval-failure-attribution&quot;&gt;12｜The genuinely hard part of agent eval: failure attribution&lt;/h2&gt;

&lt;p&gt;Suppose a financial agent gets a company’s 2026E EBITDA wrong. “Answer incorrect” is worth very little. What is worth something is knowing &lt;strong&gt;why&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Perhaps it retrieved the earnings report for the wrong year—a retrieval failure. Perhaps it found the right document but treated adjusted EBITDA as GAAP operating income—an extraction or semantic failure. Perhaps every figure was right and the formula was wrong—a calculation failure. Perhaps the analysis was correct but the citation pointed at a different filing—a citation failure. Perhaps an external tool timed out and the agent never retried—a recovery failure.&lt;/p&gt;

&lt;p&gt;So one of the most valuable artefacts in a mature agent eval is not the total score but a &lt;strong&gt;failure taxonomy&lt;/strong&gt;:&lt;/p&gt;

&lt;table&gt;
  &lt;thead&gt;
    &lt;tr&gt;
      &lt;th&gt;Failure type&lt;/th&gt;
      &lt;th&gt;Share&lt;/th&gt;
    &lt;/tr&gt;
  &lt;/thead&gt;
  &lt;tbody&gt;
    &lt;tr&gt;
      &lt;td&gt;Retrieval&lt;/td&gt;
      &lt;td&gt;27%&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Tool use&lt;/td&gt;
      &lt;td&gt;19%&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Reasoning&lt;/td&gt;
      &lt;td&gt;18%&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Data extraction&lt;/td&gt;
      &lt;td&gt;14%&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Calculation&lt;/td&gt;
      &lt;td&gt;9%&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Citation&lt;/td&gt;
      &lt;td&gt;8%&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Other&lt;/td&gt;
      &lt;td&gt;5%&lt;/td&gt;
    &lt;/tr&gt;
  &lt;/tbody&gt;
&lt;/table&gt;

&lt;p&gt;This table is usually worth far more than “overall score = 73.4,” because it answers directly what the team should fix next.&lt;/p&gt;

&lt;p&gt;If 40% of failures come from retrieval, more reasoning RL is unlikely to help. If the agent usually finds the right material but the arithmetic is unstable, what you need may be a deterministic calculator rather than a bigger model.&lt;/p&gt;

&lt;p&gt;&lt;mark&gt;The real job of eval is not to tell you how bad the model is. It is to tell you why the system is bad.&lt;/mark&gt;&lt;/p&gt;

&lt;h2 id=&quot;13why-a-single-overall-score-is-almost-never-enough&quot;&gt;13｜Why a single overall score is almost never enough&lt;/h2&gt;

&lt;p&gt;Suppose a new checkpoint moves the overall score from 82.1 to 84.0. That sounds good. Now break it apart:&lt;/p&gt;

&lt;table&gt;
  &lt;thead&gt;
    &lt;tr&gt;
      &lt;th&gt;Slice&lt;/th&gt;
      &lt;th&gt;Old&lt;/th&gt;
      &lt;th&gt;New&lt;/th&gt;
    &lt;/tr&gt;
  &lt;/thead&gt;
  &lt;tbody&gt;
    &lt;tr&gt;
      &lt;td&gt;Math&lt;/td&gt;
      &lt;td&gt;81&lt;/td&gt;
      &lt;td&gt;89&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Coding&lt;/td&gt;
      &lt;td&gt;80&lt;/td&gt;
      &lt;td&gt;87&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Writing&lt;/td&gt;
      &lt;td&gt;84&lt;/td&gt;
      &lt;td&gt;85&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Finance&lt;/td&gt;
      &lt;td&gt;86&lt;/td&gt;
      &lt;td&gt;&lt;strong&gt;77&lt;/strong&gt;&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Safety&lt;/td&gt;
      &lt;td&gt;88&lt;/td&gt;
      &lt;td&gt;&lt;strong&gt;82&lt;/strong&gt;&lt;/td&gt;
    &lt;/tr&gt;
  &lt;/tbody&gt;
&lt;/table&gt;

&lt;p&gt;If you are building a financial copilot, this is not an upgrade. It is an incident.&lt;/p&gt;

&lt;p&gt;Serious evaluation therefore requires &lt;strong&gt;slice analysis&lt;/strong&gt;, along dimensions such as task type, domain, difficulty, language, context length, tool type, risk level, user segment, and failure class.&lt;/p&gt;

&lt;p&gt;It also requires attention to statistical uncertainty. One model scoring 83% on 100 questions and another scoring 84% does not establish that the second is better, because eval is itself a sampling process. The standard error of a binomial proportion is&lt;/p&gt;

&lt;p&gt;[\mathrm{SE} \;=\; \sqrt{\frac{\hat{p}\,(1 - \hat{p})}{n}}, \qquad
\text{95\% CI} \;=\; \hat{p} \;\pm\; 1.96\,\mathrm{SE}]&lt;/p&gt;

&lt;p&gt;With (\hat p = 0.83) and (n = 100), (\mathrm{SE} \approx 3.8\%) and the interval is roughly (\pm 7.4) percentage points. The difference between 83% and 84% is almost entirely inside the noise.&lt;/p&gt;

&lt;p&gt;A better approach is paired comparison: evaluate both models on the same items and count only those where one is right and the other wrong (McNemar’s test). This removes the variance contributed by item difficulty and yields far more resolution from the same sample size.&lt;/p&gt;

&lt;p&gt;Which is why a tidy scalar on a leaderboard is, in real model development, only the entry point. What matters is:&lt;/p&gt;

&lt;blockquote&gt;
  &lt;p&gt;&lt;strong&gt;Where did we improve, where did we regress, and why?&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2 id=&quot;14offline-eval-online-eval-and-the-real-worlds-final-vote&quot;&gt;14｜Offline eval, online eval, and the real world’s final vote&lt;/h2&gt;

&lt;p&gt;Even the best golden set has an intrinsic problem: it is a proxy for the real world. Whether a product succeeds is not decided by MMLU, SWE-bench, or an internal judge score. It is decided by real users.&lt;/p&gt;

&lt;p&gt;So eval splits into two worlds.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Offline eval&lt;/strong&gt; serves development. It is cheap, fast, repeatable, supports regression testing, and catches obvious problems before release.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Online eval&lt;/strong&gt; looks at real production outcomes: task completion, user preference, regeneration rate, escalation rate, retention, conversion, latency, token cost, cost per successful task.&lt;/p&gt;

&lt;p&gt;That last metric deserves to be written out, because it often reflects a system’s true economics better than accuracy does:&lt;/p&gt;

&lt;p&gt;[\text{cost per successful task} \;=\; \frac{\mathbb{E}\big[\text{cost per attempt}\big]}{\Pr\big[\text{success}\big]}]&lt;/p&gt;

&lt;p&gt;A change that lifts success from 60% to 75% can be cheaper overall even if each call costs more.&lt;/p&gt;

&lt;p&gt;But there is a trap here too: &lt;strong&gt;user preference is not truth&lt;/strong&gt;. A model that is confident, well-spoken, and agrees with the user can score very well on immediate preference while being poorly calibrated and factually weak. The real product objective is usually multi-objective, and stating it with constraints is more honest:&lt;/p&gt;

&lt;p&gt;[\max_{\text{system}} \;\; \mathbb{E}\big[\text{task success}\big]
\quad \text{s.t.} \quad
\text{factuality} \ge \tau_f,\;\;
\text{harm rate} \le \tau_s,\;\;
\mathbb{E}[\text{cost}] \le c,\;\;
p_{95}(\text{latency}) \le \ell]&lt;/p&gt;

&lt;p&gt;Putting safety and factuality into the constraints rather than into a weighted sum is the point: they should not be purchasable with gains in success rate.&lt;/p&gt;

&lt;p&gt;Which is why there is no such thing as a &lt;em&gt;universal LLM score&lt;/em&gt;. &lt;mark&gt;Capability is not a scalar; it is a vector.&lt;/mark&gt; “Which model is best” is an incomplete question. The real one is always:&lt;/p&gt;

&lt;blockquote&gt;
  &lt;p&gt;&lt;strong&gt;Best for what?&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2 id=&quot;15so-how-should-you-actually-build-an-eval-system&quot;&gt;15｜So how should you actually build an eval system?&lt;/h2&gt;

&lt;p&gt;Starting an LLM or agent product from scratch today, I would not begin by asking which benchmarks the industry uses. I would work through the following.&lt;/p&gt;

&lt;ol&gt;
  &lt;li&gt;&lt;strong&gt;Collect real tasks&lt;/strong&gt;, rather than having a PM invent prompts in a meeting room. You need to know the actual workload distribution.&lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Build a taxonomy&lt;/strong&gt;: task type × difficulty × domain × risk × tool.&lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Extract a high-quality golden set.&lt;/strong&gt; It need not be large, but it must be representative, and it must include edge cases and historical production failures.&lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Write a rubric for each task type.&lt;/strong&gt; What counts as fully correct, as partial success, as critical failure; whether abstention is allowed; whether citations are required.&lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Never reach for an LLM judge where deterministic grading is possible.&lt;/strong&gt; Run the tests for code, compute the number directly, check the final environment state for tool tasks.&lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;For tasks that cannot be executed&lt;/strong&gt;—writing, analysis, open-ended QA—introduce a criteria-based LLM judge, and calibrate it against expert human labels.&lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Report more than an overall score.&lt;/strong&gt; Maintain slices and a failure taxonomy.&lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Run the regression suite on every change&lt;/strong&gt;—model, prompt, retrieval, tool, system prompt. All of it counts.&lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Feed new production failures back into the eval set continuously.&lt;/strong&gt;&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The development process then becomes a closed loop:&lt;/p&gt;

&lt;blockquote&gt;
  &lt;p&gt;real workload → golden set → rubric → grader → slice and failure analysis → locate the bottleneck → change model / prompt / retrieval / tool → regression → ship → production failures → back into the golden set&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;This is &lt;strong&gt;eval-driven development&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;It resembles test-driven development in ordinary software engineering, but it is considerably harder, because conventional software testing targets a deterministic system while an LLM is stochastic, open-ended, and increasingly interacts with an environment of its own accord.&lt;/p&gt;

&lt;h2 id=&quot;16why-eval-ends-up-being-the-core-infrastructure-of-ai-development&quot;&gt;16｜Why eval ends up being the core infrastructure of AI development&lt;/h2&gt;

&lt;p&gt;We can now read the history again.&lt;/p&gt;

&lt;p&gt;In the MMLU era, evaluation mostly meant &lt;strong&gt;setting the model an exam&lt;/strong&gt;. HumanEval taught us &lt;strong&gt;not to look at text but to execute what the model produced&lt;/strong&gt;. SWE-bench taught us &lt;strong&gt;not to test isolated tasks but real workflows&lt;/strong&gt;. Chatbot Arena taught us that &lt;strong&gt;some qualities have no unique answer and can only be learned from human preference&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;RLHF trained human preference into a reward model, and &lt;strong&gt;evaluation became a training objective&lt;/strong&gt;. RLAIF let AI produce the preferences, and &lt;strong&gt;the evaluator began to scale too&lt;/strong&gt;. Process reward showed that for complex reasoning we may need to supervise the path and not only the outcome. Best-of-N moved the evaluator to inference time: the model generates many possibilities and the verifier decides which survives. And agent benchmarks pushed the whole question up to the level of the environment—we no longer evaluate what the model said, but what the system actually accomplished.&lt;/p&gt;

&lt;p&gt;To read evaluation today as “running a few benchmarks after training” is to underrate it entirely. In a growing number of modern AI systems the evaluator plays at least four roles at once: &lt;mark&gt;it measures the model and it shapes the model; it defines the reward and it guides inference-time search.&lt;/mark&gt; It tells you which model is stronger, and it tells you why your system failed.&lt;/p&gt;

&lt;p&gt;And as foundation models themselves become easier to obtain—APIs to call, open-weight models to download, increasingly standardised fine-tuning pipelines—what is genuinely hard to copy may turn out to be domain data, expert feedback, production failure history, and evaluation infrastructure accumulated over years.&lt;/p&gt;

&lt;p&gt;Because the most dangerous thing in AI development has never been that the model failed to improve. It is:&lt;/p&gt;

&lt;blockquote&gt;
  &lt;p&gt;&lt;strong&gt;that you believed it had.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;A model gains five points on a leaderboard and regresses on your core user task. A reward model score climbs steadily while the policy merely gets better at hacking it. An agent demo looks extraordinary and then falls apart across 1,000 production cases.&lt;/p&gt;

&lt;p&gt;Without good eval, you do not even have the language to describe these problems.&lt;/p&gt;

&lt;p&gt;So the question that matters most for model teams may not stay &lt;em&gt;How do we train a smarter model?&lt;/em&gt; It becomes, increasingly:&lt;/p&gt;

&lt;blockquote&gt;
  &lt;p&gt;&lt;strong&gt;What does “smarter” actually mean, and how do we know when we get there?&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That is evaluation. On the surface it looks like scoring an AI. In practice, it is where we decide what we want the AI to become.&lt;/p&gt;
</content>
  </entry>
  
  <entry>
    <title>观拾麦</title>
    <link href="https://mochiaochen.github.io/writing/2026/08/gleaning-wheat/" rel="alternate" type="text/html"/>
    <published>2026-08-28T21:00:00+08:00</published>
    <updated>2026-08-28T21:00:00+08:00</updated>
    <id>https://mochiaochen.github.io/writing/2026/08/gleaning-wheat</id>
    <content type="html" xml:base="https://mochiaochen.github.io/writing/2026/08/gleaning-wheat/">&lt;figure&gt;
  &lt;p&gt;&lt;img src=&quot;/assets/images/millet-gleaners.jpg&quot; alt=&quot;油画：暮色下的麦田里，三名妇人弯腰俯身，拾取收割后遗落的麦穗，远处是堆垛与忙碌的收割队伍&quot; width=&quot;1600&quot; height=&quot;1197&quot; /&gt;&lt;/p&gt;
  &lt;figcaption&gt;
    &lt;p&gt;让-弗朗索瓦 · 米勒（Jean-François Millet）《拾穗者》（Des glaneuses），1857 年，布面油画，巴黎奥赛博物馆藏。公有领域，图像来自 Wikimedia Commons。&amp;lt;/figcaption&amp;gt;
&amp;lt;/figure&amp;gt;&lt;/p&gt;

    &lt;p class=&quot;verse&quot;&gt;秋气满平冥&lt;sup id=&quot;fnref:2&quot; role=&quot;doc-noteref&quot;&gt;&lt;a href=&quot;#fn:2&quot; class=&quot;footnote&quot; rel=&quot;footnote&quot;&gt;1&lt;/a&gt;&lt;/sup&gt;，寒烟起晚汀&lt;sup id=&quot;fnref:3&quot; role=&quot;doc-noteref&quot;&gt;&lt;a href=&quot;#fn:3&quot; class=&quot;footnote&quot; rel=&quot;footnote&quot;&gt;2&lt;/a&gt;&lt;/sup&gt;。&lt;br /&gt;
尘随衣上起，汗入日中清&lt;sup id=&quot;fnref:4&quot; role=&quot;doc-noteref&quot;&gt;&lt;a href=&quot;#fn:4&quot; class=&quot;footnote&quot; rel=&quot;footnote&quot;&gt;3&lt;/a&gt;&lt;/sup&gt;。&lt;br /&gt;
命薄&lt;sup id=&quot;fnref:5&quot; role=&quot;doc-noteref&quot;&gt;&lt;a href=&quot;#fn:5&quot; class=&quot;footnote&quot; rel=&quot;footnote&quot;&gt;4&lt;/a&gt;&lt;/sup&gt;年年苦，心慈步步轻。&lt;br /&gt;
长嗟天地厚，不救一身贫&lt;sup id=&quot;fnref:6&quot; role=&quot;doc-noteref&quot;&gt;&lt;a href=&quot;#fn:6&quot; class=&quot;footnote&quot; rel=&quot;footnote&quot;&gt;5&lt;/a&gt;&lt;/sup&gt;。&lt;/p&gt;

    &lt;p class=&quot;postscript&quot;&gt;尝赴民大&lt;sup id=&quot;fnref:7&quot; role=&quot;doc-noteref&quot;&gt;&lt;a href=&quot;#fn:7&quot; class=&quot;footnote&quot; rel=&quot;footnote&quot;&gt;6&lt;/a&gt;&lt;/sup&gt;就试北京市人文知识竞赛，题命拾麦&lt;sup id=&quot;fnref:1&quot; role=&quot;doc-noteref&quot;&gt;&lt;a href=&quot;#fn:1&quot; class=&quot;footnote&quot; rel=&quot;footnote&quot;&gt;7&lt;/a&gt;&lt;/sup&gt;图，作五律一首，限清韵&lt;sup id=&quot;fnref:8&quot; role=&quot;doc-noteref&quot;&gt;&lt;a href=&quot;#fn:8&quot; class=&quot;footnote&quot; rel=&quot;footnote&quot;&gt;8&lt;/a&gt;&lt;/sup&gt;。时急成之，未遑雕琢&lt;sup id=&quot;fnref:9&quot; role=&quot;doc-noteref&quot;&gt;&lt;a href=&quot;#fn:9&quot; class=&quot;footnote&quot; rel=&quot;footnote&quot;&gt;9&lt;/a&gt;&lt;/sup&gt;。今展卷重温，稍加润色，庶几&lt;sup id=&quot;fnref:10&quot; role=&quot;doc-noteref&quot;&gt;&lt;a href=&quot;#fn:10&quot; class=&quot;footnote&quot; rel=&quot;footnote&quot;&gt;10&lt;/a&gt;&lt;/sup&gt;不负当日之思。&lt;/p&gt;

    &lt;hr /&gt;

    &lt;div class=&quot;footnotes&quot; role=&quot;doc-endnotes&quot;&gt;
      &lt;ol&gt;
    &lt;li id=&quot;fn:2&quot; role=&quot;doc-endnote&quot;&gt;
      &lt;p&gt;&lt;strong&gt;平冥&lt;/strong&gt;：平旷苍茫的远天。冥，幽暗、深远。谓秋气弥漫，直至天际。 &lt;a href=&quot;#fnref:2&quot; class=&quot;reversefootnote&quot; role=&quot;doc-backlink&quot;&gt;&amp;#8617;&lt;/a&gt;&lt;/p&gt;
    &lt;/li&gt;
    &lt;li id=&quot;fn:3&quot; role=&quot;doc-endnote&quot;&gt;
      &lt;p&gt;&lt;strong&gt;汀（tīng）&lt;/strong&gt;：水边平地或水中小洲。语出谢朓《之宣城郡出新林浦向板桥》：「天际识归舟，云中辨江树。」一类江汀晚景。此句写日暮水气升腾，寒意先至。 &lt;a href=&quot;#fnref:3&quot; class=&quot;reversefootnote&quot; role=&quot;doc-backlink&quot;&gt;&amp;#8617;&lt;/a&gt;&lt;/p&gt;
    &lt;/li&gt;
    &lt;li id=&quot;fn:4&quot; role=&quot;doc-endnote&quot;&gt;
      &lt;p&gt;&lt;strong&gt;汗入日中清&lt;/strong&gt;：日中，正午。谓烈日之下汗流浃背，反而洗得人身心澄澈。与上句「尘随衣上起」相对：一浊一清，一起一落。 &lt;a href=&quot;#fnref:4&quot; class=&quot;reversefootnote&quot; role=&quot;doc-backlink&quot;&gt;&amp;#8617;&lt;/a&gt;&lt;/p&gt;
    &lt;/li&gt;
    &lt;li id=&quot;fn:5&quot; role=&quot;doc-endnote&quot;&gt;
      &lt;p&gt;&lt;strong&gt;命薄&lt;/strong&gt;：命运浅薄，谓身世寒微、生计艰难。 &lt;a href=&quot;#fnref:5&quot; class=&quot;reversefootnote&quot; role=&quot;doc-backlink&quot;&gt;&amp;#8617;&lt;/a&gt;&lt;/p&gt;
    &lt;/li&gt;
    &lt;li id=&quot;fn:6&quot; role=&quot;doc-endnote&quot;&gt;
      &lt;p&gt;&lt;strong&gt;长嗟天地厚，不救一身贫&lt;/strong&gt;：天地覆载之德不可谓不厚，却救不得一人之贫。语意近于杜甫「安得广厦千万间，大庇天下寒士俱欢颜」的悲悯，而出以诘问，语更沉痛。 &lt;a href=&quot;#fnref:6&quot; class=&quot;reversefootnote&quot; role=&quot;doc-backlink&quot;&gt;&amp;#8617;&lt;/a&gt;&lt;/p&gt;
    &lt;/li&gt;
    &lt;li id=&quot;fn:7&quot; role=&quot;doc-endnote&quot;&gt;
      &lt;p&gt;&lt;strong&gt;民大&lt;/strong&gt;：中央民族大学，在北京市海淀区中关村南大街。 &lt;a href=&quot;#fnref:7&quot; class=&quot;reversefootnote&quot; role=&quot;doc-backlink&quot;&gt;&amp;#8617;&lt;/a&gt;&lt;/p&gt;
    &lt;/li&gt;
    &lt;li id=&quot;fn:1&quot; role=&quot;doc-endnote&quot;&gt;
      &lt;p&gt;&lt;strong&gt;拾麦&lt;/strong&gt;：收割之后，拾取遗落于田间的麦穗，旧时贫者藉以糊口。唐 · 白居易《观刈麦》写此最切：「复有贫妇人，抱子在其旁。右手秉遗穗，左臂悬敝筐。」诗题「观拾麦」，正从「观刈麦」化出。 &lt;a href=&quot;#fnref:1&quot; class=&quot;reversefootnote&quot; role=&quot;doc-backlink&quot;&gt;&amp;#8617;&lt;/a&gt;&lt;/p&gt;
    &lt;/li&gt;
    &lt;li id=&quot;fn:8&quot; role=&quot;doc-endnote&quot;&gt;
      &lt;p&gt;&lt;strong&gt;限清韵&lt;/strong&gt;：旧时试帖诗与赛诗多有限韵之例，即指定所押韵部，不得出韵。「限清韵」谓限用「清」字所属之韵部。 &lt;a href=&quot;#fnref:8&quot; class=&quot;reversefootnote&quot; role=&quot;doc-backlink&quot;&gt;&amp;#8617;&lt;/a&gt;&lt;/p&gt;
    &lt;/li&gt;
    &lt;li id=&quot;fn:9&quot; role=&quot;doc-endnote&quot;&gt;
      &lt;p&gt;&lt;strong&gt;未遑（huáng）雕琢&lt;/strong&gt;：来不及推敲修饰。遑，闲暇。 &lt;a href=&quot;#fnref:9&quot; class=&quot;reversefootnote&quot; role=&quot;doc-backlink&quot;&gt;&amp;#8617;&lt;/a&gt;&lt;/p&gt;
    &lt;/li&gt;
    &lt;li id=&quot;fn:10&quot; role=&quot;doc-endnote&quot;&gt;
      &lt;p&gt;&lt;strong&gt;庶几（shù jī）&lt;/strong&gt;：差不多、但愿，表希冀之词。《孟子 · 梁惠王下》：「王之好乐甚，则齐国其庶几乎。」 &lt;a href=&quot;#fnref:10&quot; class=&quot;reversefootnote&quot; role=&quot;doc-backlink&quot;&gt;&amp;#8617;&lt;/a&gt;&lt;/p&gt;
    &lt;/li&gt;
  &lt;/ol&gt;
    &lt;/div&gt;
  &lt;/figcaption&gt;
&lt;/figure&gt;
</content>
  </entry>
  
  <entry>
    <title>Watching the Gleaners</title>
    <link href="https://mochiaochen.github.io/en/writing/2026/08/gleaning-wheat/" rel="alternate" type="text/html"/>
    <published>2026-08-28T21:00:00+08:00</published>
    <updated>2026-08-28T21:00:00+08:00</updated>
    <id>https://mochiaochen.github.io/en/writing/2026/08/gleaning-wheat-en</id>
    <content type="html" xml:base="https://mochiaochen.github.io/en/writing/2026/08/gleaning-wheat/">&lt;figure&gt;
  &lt;p&gt;&lt;img src=&quot;/assets/images/millet-gleaners.jpg&quot; alt=&quot;Oil painting: three women stoop in a stubble field at the end of the day, gathering fallen ears of wheat, with stacks and a busy harvest crew behind them&quot; width=&quot;1600&quot; height=&quot;1197&quot; /&gt;&lt;/p&gt;
  &lt;figcaption&gt;
    &lt;p&gt;Jean-François Millet, &lt;em&gt;Des glaneuses&lt;/em&gt; (The Gleaners), 1857. Oil on canvas, Musée d’Orsay, Paris. Public domain, via Wikimedia Commons.&amp;lt;/figcaption&amp;gt;
&amp;lt;/figure&amp;gt;&lt;/p&gt;

    &lt;p class=&quot;verse verse--flush&quot;&gt;Autumn air fills the flat and darkening sky;&lt;sup id=&quot;fnref:2-en&quot; role=&quot;doc-noteref&quot;&gt;&lt;a href=&quot;#fn:2-en&quot; class=&quot;footnote&quot; rel=&quot;footnote&quot;&gt;1&lt;/a&gt;&lt;/sup&gt;&lt;br /&gt;
cold mist rises from the evening shoal.&lt;sup id=&quot;fnref:3-en&quot; role=&quot;doc-noteref&quot;&gt;&lt;a href=&quot;#fn:3-en&quot; class=&quot;footnote&quot; rel=&quot;footnote&quot;&gt;2&lt;/a&gt;&lt;/sup&gt;&lt;br /&gt;
Dust lifts with every fold of cloth;&lt;br /&gt;
sweat runs clear in the noonday sun.&lt;sup id=&quot;fnref:4-en&quot; role=&quot;doc-noteref&quot;&gt;&lt;a href=&quot;#fn:4-en&quot; class=&quot;footnote&quot; rel=&quot;footnote&quot;&gt;3&lt;/a&gt;&lt;/sup&gt;&lt;br /&gt;
A thin fate:&lt;sup id=&quot;fnref:5-en&quot; role=&quot;doc-noteref&quot;&gt;&lt;a href=&quot;#fn:5-en&quot; class=&quot;footnote&quot; rel=&quot;footnote&quot;&gt;4&lt;/a&gt;&lt;/sup&gt; bitter, year upon year.&lt;br /&gt;
A kind heart: light with every step.&lt;br /&gt;
Long I sigh that heaven and earth are so generous,&lt;br /&gt;
and cannot save one body from its poverty.&lt;sup id=&quot;fnref:6-en&quot; role=&quot;doc-noteref&quot;&gt;&lt;a href=&quot;#fn:6-en&quot; class=&quot;footnote&quot; rel=&quot;footnote&quot;&gt;5&lt;/a&gt;&lt;/sup&gt;&lt;/p&gt;

    &lt;p class=&quot;postscript&quot;&gt;I once travelled to Minzu University of China&lt;sup id=&quot;fnref:7-en&quot; role=&quot;doc-noteref&quot;&gt;&lt;a href=&quot;#fn:7-en&quot; class=&quot;footnote&quot; rel=&quot;footnote&quot;&gt;6&lt;/a&gt;&lt;/sup&gt; to sit the Beijing municipal humanities competition. The set subject was a painting of gleaners,&lt;sup id=&quot;fnref:1-en&quot; role=&quot;doc-noteref&quot;&gt;&lt;a href=&quot;#fn:1-en&quot; class=&quot;footnote&quot; rel=&quot;footnote&quot;&gt;7&lt;/a&gt;&lt;/sup&gt; and I wrote a poem in five-character regulated verse with the rhyme fixed in advance.&lt;sup id=&quot;fnref:8-en&quot; role=&quot;doc-noteref&quot;&gt;&lt;a href=&quot;#fn:8-en&quot; class=&quot;footnote&quot; rel=&quot;footnote&quot;&gt;8&lt;/a&gt;&lt;/sup&gt; It was finished in haste, with no time to polish it.&lt;sup id=&quot;fnref:9-en&quot; role=&quot;doc-noteref&quot;&gt;&lt;a href=&quot;#fn:9-en&quot; class=&quot;footnote&quot; rel=&quot;footnote&quot;&gt;9&lt;/a&gt;&lt;/sup&gt; Opening the sheet again today, I have touched it up a little, in the hope&lt;sup id=&quot;fnref:10-en&quot; role=&quot;doc-noteref&quot;&gt;&lt;a href=&quot;#fn:10-en&quot; class=&quot;footnote&quot; rel=&quot;footnote&quot;&gt;10&lt;/a&gt;&lt;/sup&gt; of not falling short of what I meant that day.&lt;/p&gt;

    &lt;hr /&gt;

    &lt;div class=&quot;footnotes&quot; role=&quot;doc-endnotes&quot;&gt;
      &lt;ol&gt;
    &lt;li id=&quot;fn:2-en&quot; role=&quot;doc-endnote&quot;&gt;
      &lt;p&gt;&lt;strong&gt;The flat and darkening sky:&lt;/strong&gt; the original compounds “level” with a word meaning dim, deep, remote—the whole expanse of sky above open country, filling with autumn. &lt;a href=&quot;#fnref:2-en&quot; class=&quot;reversefootnote&quot; role=&quot;doc-backlink&quot;&gt;&amp;#8617;&lt;/a&gt;&lt;/p&gt;
    &lt;/li&gt;
    &lt;li id=&quot;fn:3-en&quot; role=&quot;doc-endnote&quot;&gt;
      &lt;p&gt;&lt;strong&gt;Shoal:&lt;/strong&gt; the flat ground at the water’s edge, or a small islet in a stream. The line sets the hour: the mist rising off the water at dusk, cold arriving before the dark. &lt;a href=&quot;#fnref:3-en&quot; class=&quot;reversefootnote&quot; role=&quot;doc-backlink&quot;&gt;&amp;#8617;&lt;/a&gt;&lt;/p&gt;
    &lt;/li&gt;
    &lt;li id=&quot;fn:4-en&quot; role=&quot;doc-endnote&quot;&gt;
      &lt;p&gt;The couplet is built on a reversal: dust rises and clouds the clothes; sweat pours in the midday heat and leaves the body clear. One line turbid, one clear; one rising, one washing away. &lt;a href=&quot;#fnref:4-en&quot; class=&quot;reversefootnote&quot; role=&quot;doc-backlink&quot;&gt;&amp;#8617;&lt;/a&gt;&lt;/p&gt;
    &lt;/li&gt;
    &lt;li id=&quot;fn:5-en&quot; role=&quot;doc-endnote&quot;&gt;
      &lt;p&gt;&lt;strong&gt;A thin fate:&lt;/strong&gt; a humble station and a hard livelihood. &lt;a href=&quot;#fnref:5-en&quot; class=&quot;reversefootnote&quot; role=&quot;doc-backlink&quot;&gt;&amp;#8617;&lt;/a&gt;&lt;/p&gt;
    &lt;/li&gt;
    &lt;li id=&quot;fn:6-en&quot; role=&quot;doc-endnote&quot;&gt;
      &lt;p&gt;The closing couplet: the covering virtue of heaven and earth could hardly be called anything but generous, and yet it cannot lift a single person out of poverty. The compassion is close to Du Fu’s wish for “a mansion of ten thousand rooms, to shelter every poor scholar beneath heaven,” but put as a question, and sharper for it. &lt;a href=&quot;#fnref:6-en&quot; class=&quot;reversefootnote&quot; role=&quot;doc-backlink&quot;&gt;&amp;#8617;&lt;/a&gt;&lt;/p&gt;
    &lt;/li&gt;
    &lt;li id=&quot;fn:7-en&quot; role=&quot;doc-endnote&quot;&gt;
      &lt;p&gt;&lt;strong&gt;Minzu University of China:&lt;/strong&gt; on South Zhongguancun Street in the Haidian district of Beijing. &lt;a href=&quot;#fnref:7-en&quot; class=&quot;reversefootnote&quot; role=&quot;doc-backlink&quot;&gt;&amp;#8617;&lt;/a&gt;&lt;/p&gt;
    &lt;/li&gt;
    &lt;li id=&quot;fn:1-en&quot; role=&quot;doc-endnote&quot;&gt;
      &lt;p&gt;&lt;strong&gt;Gleaning:&lt;/strong&gt; gathering the ears of grain left behind after the harvest, on which the poor once depended. The classic treatment is Bai Juyi’s “Watching the Wheat Harvest” (Tang dynasty): “And there is a poor woman too, her child held at her side—in her right hand the fallen ears, on her left arm a broken basket.” The title of this poem, “Watching the Gleaners,” is formed directly on Bai Juyi’s. &lt;a href=&quot;#fnref:1-en&quot; class=&quot;reversefootnote&quot; role=&quot;doc-backlink&quot;&gt;&amp;#8617;&lt;/a&gt;&lt;/p&gt;
    &lt;/li&gt;
    &lt;li id=&quot;fn:8-en&quot; role=&quot;doc-endnote&quot;&gt;
      &lt;p&gt;&lt;strong&gt;A fixed rhyme:&lt;/strong&gt; examination poems and poetry contests traditionally specified the rhyme group a poem had to use throughout, with departures counted as faults. Here the rhyme group was the one containing the character &lt;em&gt;qing&lt;/em&gt;, “clear.” &lt;a href=&quot;#fnref:8-en&quot; class=&quot;reversefootnote&quot; role=&quot;doc-backlink&quot;&gt;&amp;#8617;&lt;/a&gt;&lt;/p&gt;
    &lt;/li&gt;
    &lt;li id=&quot;fn:9-en&quot; role=&quot;doc-endnote&quot;&gt;
      &lt;p&gt;The word rendered “no time” literally means leisure: there was none. &lt;a href=&quot;#fnref:9-en&quot; class=&quot;reversefootnote&quot; role=&quot;doc-backlink&quot;&gt;&amp;#8617;&lt;/a&gt;&lt;/p&gt;
    &lt;/li&gt;
    &lt;li id=&quot;fn:10-en&quot; role=&quot;doc-endnote&quot;&gt;
      &lt;p&gt;&lt;strong&gt;In the hope:&lt;/strong&gt; a classical particle of tentative wishing, familiar from Mencius: “If Your Majesty’s love of music is as great as this, the state of Qi is perhaps not far from good order.” &lt;a href=&quot;#fnref:10-en&quot; class=&quot;reversefootnote&quot; role=&quot;doc-backlink&quot;&gt;&amp;#8617;&lt;/a&gt;&lt;/p&gt;
    &lt;/li&gt;
  &lt;/ol&gt;
    &lt;/div&gt;
  &lt;/figcaption&gt;
&lt;/figure&gt;
</content>
  </entry>
  
  <entry>
    <title>新生产力，需要新的生产关系</title>
    <link href="https://mochiaochen.github.io/writing/2026/07/new-productive-forces-new-relations-of-production/" rel="alternate" type="text/html"/>
    <published>2026-07-24T10:00:00+08:00</published>
    <updated>2026-07-24T10:00:00+08:00</updated>
    <id>https://mochiaochen.github.io/writing/2026/07/new-productive-forces-new-relations-of-production</id>
    <content type="html" xml:base="https://mochiaochen.github.io/writing/2026/07/new-productive-forces-new-relations-of-production/">&lt;p&gt;&lt;mark&gt;任何一次生产力的跃迁，最终都要靠新的生产关系来承接。&lt;/mark&gt;否则，新的能力会被旧的结构压住，无法释放。&lt;/p&gt;

&lt;p&gt;SpaceX 为什么成功？技术并不是 NASA 没有——NASA 有。真正的原因是，NASA 是旧生产关系的产物。官僚体制、分包商网络、国会预算周期，这些生产关系是为上一代生产力设计的。马斯克用全新的生产关系——第一性原理开发、垂直整合、快速迭代、研究工程一体——把同样的物理学做出了完全不同的结果。&lt;/p&gt;

&lt;p&gt;所以问题来了：&lt;strong&gt;Agentic AI 作为新生产力，它本质上改变了什么？&lt;/strong&gt;&lt;/p&gt;

&lt;h2 id=&quot;一劳动的基本单位变了&quot;&gt;一、劳动的基本单位变了&lt;/h2&gt;

&lt;p&gt;在农业时代，生产的基本单位是&lt;strong&gt;人+土地&lt;/strong&gt;。在工业时代，是&lt;strong&gt;人+机器&lt;/strong&gt;。在数字化时代，是&lt;strong&gt;工程师+代码+数据&lt;/strong&gt;。&lt;/p&gt;

&lt;p&gt;今天，当硅基智能能够自主规划、自主执行、持续产出的时候，生产的基本单位开始变成：&lt;/p&gt;

&lt;blockquote&gt;
  &lt;p&gt;&lt;strong&gt;人类判断力 + AI Agent 系统&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;这是历史上第一次，认知劳动本身可以被自动化执行。以前的工业革命替代的是体力劳动，今天被替代的是执行层的认知劳动：写代码的执行、写报告的执行、做分析的执行，都开始由硅基智能承接。&lt;/p&gt;

&lt;p&gt;这意味着，人的价值将越来越集中在两种能力上：&lt;/p&gt;

&lt;ol&gt;
  &lt;li&gt;&lt;strong&gt;提出正确问题的能力；&lt;/strong&gt;&lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;判断什么有价值的品味。&lt;/strong&gt;&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;mark&gt;执行变得廉价之后，问题与品味就会变得昂贵。&lt;/mark&gt;&lt;/p&gt;

&lt;h2 id=&quot;二组织规模走向哑铃结构&quot;&gt;二、组织规模走向「哑铃结构」&lt;/h2&gt;

&lt;p&gt;组织的最优规模，正在两个方向上同时发生剧变。&lt;/p&gt;

&lt;p&gt;一方面，对于前沿研究，规模必须越来越大。训练一代新模型，可能需要数百亿美元的算力投入。这推动 &lt;strong&gt;NeoLab&lt;/strong&gt; 作为新的组织形态出现——它必须同时承载科学风险、工程风险、市场风险和资本风险，&lt;strong&gt;小而精，却又资本密集&lt;/strong&gt;。&lt;/p&gt;

&lt;p&gt;另一方面，在应用和价值创造层，最优规模正在急剧缩小。&lt;strong&gt;FDE&lt;/strong&gt; 是一个信号，&lt;strong&gt;OPC（One Person Company，一人公司）&lt;/strong&gt;则是更极端的形态。过去需要十人团队才能交付的产品，今天一个高判断力的人配合 AI Agent 系统就可以完成。Major Shlomo 用一个人、六个月，做到了 2500 万美元的收购估值。&lt;/p&gt;

&lt;p&gt;所以，新的生产关系在组织层面表现为一个「哑铃结构」：&lt;/p&gt;

&lt;ul&gt;
  &lt;li&gt;顶部是少数超大规模的 &lt;strong&gt;NeoLab&lt;/strong&gt;，集中研究密度、算力资源和顶尖人才；&lt;/li&gt;
  &lt;li&gt;底部是大量小型 &lt;strong&gt;OPC 与 FDE&lt;/strong&gt;，高度灵活地把顶部的能力接入具体产业场景；&lt;/li&gt;
  &lt;li&gt;中间则是传统企业——尤其是依靠人头规模维持存在的组织——它们会承受最大的压力。&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;mark&gt;未来最稀缺的未必是规模，而是高密度研究与高密度判断；夹在两者之间的低密度组织最危险。&lt;/mark&gt;&lt;/p&gt;

&lt;h2 id=&quot;三资本与研究从融资走向共同生产&quot;&gt;三、资本与研究从「融资」走向「共同生产」&lt;/h2&gt;

&lt;p&gt;过去的 VC 模式是：&lt;strong&gt;我给你钱，你在七年内给我回报。&lt;/strong&gt;这个模式的设计假设是，创业是一个快速验证市场的过程，周期短，风险可分散。&lt;/p&gt;

&lt;p&gt;但研究型创业所承担的资本风险是另一个量级。训练前沿模型的算力成本、脑机接口进入临床验证的时间成本、量子计算硬件从实验室走向产业化的路径长度——这些都远超传统 VC 周期的边界。&lt;/p&gt;

&lt;p&gt;因此，新的生产关系要求出现&lt;strong&gt;耐心资本&lt;/strong&gt;：主权财富基金、大学捐赠基金，以及愿意以十年甚至二十年为单位配置资源的长期资本。这类资本进入生产关系，不只是在融资，而是在与研究型创业者共同承担生产过程本身的风险。&lt;/p&gt;

&lt;p&gt;Anthropic 与亚马逊的关系、OpenAI 与微软的关系，已经不再是传统意义上的投资人与被投资人，&lt;strong&gt;更接近共同生产的合作结构&lt;/strong&gt;。&lt;/p&gt;

&lt;h2 id=&quot;四算力基础设施正在国家化&quot;&gt;四、算力基础设施正在国家化&lt;/h2&gt;

&lt;p&gt;互联网时代的生产关系，在很大程度上是跨国的、去地域的。一家硅谷公司可以为全球用户提供服务，代码在云端，数据仿佛没有边界。&lt;/p&gt;

&lt;p&gt;算力完全不同。&lt;strong&gt;算力是物理的&lt;/strong&gt;：它依赖能源，依赖特殊制程的芯片，也依赖地理环境。这使国家重新成为生产关系中的关键角色。&lt;/p&gt;

&lt;p&gt;美国的 StarGate、中国的实训场，本质上都是国家在直接组织算力基础设施的生产关系。这是继工业时代的铁路与电网之后，国家第一次以这种力度重新介入生产基础设施。&lt;/p&gt;

&lt;p&gt;新的生产关系因此多了一个维度：它不只是市场中企业与企业、资本与劳动之间的关系，也包括&lt;strong&gt;大国之间围绕核心生产要素控制权的博弈&lt;/strong&gt;。&lt;/p&gt;

&lt;p&gt;&lt;mark&gt;当算力成为基础设施，技术竞争也就同时成为能源竞争、制造竞争与国家能力的竞争。&lt;/mark&gt;&lt;/p&gt;

&lt;h2 id=&quot;五价值分配出现最大的分叉&quot;&gt;五、价值分配出现最大的分叉&lt;/h2&gt;

&lt;p&gt;旧的价值分配逻辑具有一定的分散性：企业创造价值，雇员分享价值，市场再通过竞争使价值向用户扩散。&lt;/p&gt;

&lt;p&gt;新的生产力结构会造成一种极端的&lt;strong&gt;集中—扩散二元结构&lt;/strong&gt;。&lt;/p&gt;

&lt;p&gt;在生产力层——算力、模型、数据——价值极度集中。少数掌握算力和前沿模型的组织，拥有接近垄断的优势。&lt;/p&gt;

&lt;p&gt;但在价值应用层，AI 的普惠性又会让原本无力进入某些领域的人获得能力。FDE 可以让一名工程师深入医疗、法律、材料等行业，完成过去只有顶尖专家团队才能做到的事。&lt;/p&gt;

&lt;p&gt;这是一个矛盾的结果：&lt;/p&gt;

&lt;blockquote&gt;
  &lt;p&gt;&lt;strong&gt;生产资料的集中，同时伴随着生产能力的扩散。&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;历史上每一次大的生产力革命都有这个特征。蒸汽机让少数人极度富有，同时也把工厂带到了每一座城市。今天同样如此。&lt;/p&gt;

&lt;h2 id=&quot;六制度选择决定分叉的方向&quot;&gt;六、制度选择决定分叉的方向&lt;/h2&gt;

&lt;p&gt;这个分叉最终走向哪里，取决于我们能不能在制度上做出正确的选择：&lt;/p&gt;

&lt;ul&gt;
  &lt;li&gt;耐心资本能不能支持真正的研究；&lt;/li&gt;
  &lt;li&gt;普惠性的基础设施能不能被建立；&lt;/li&gt;
  &lt;li&gt;FDE 的能力能不能被更多人获得。&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;这些问题已经超出了任何个体或单一组织能够独立回答的范围。但它们正在变得越来越重要。&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Agentic AI 带来的并不只是一次工具升级。&lt;/strong&gt;它正在重写劳动、组织、资本、国家与分配之间的关系。新生产力已经出现，真正悬而未决的问题是：我们能否及时创造出承接它的新生产关系。&lt;/p&gt;

&lt;p&gt;&lt;mark&gt;如果不能，新能力只会强化旧结构；如果能够，它才可能真正转化为一个时代的生产力。&lt;/mark&gt;&lt;/p&gt;
</content>
  </entry>
  
  <entry>
    <title>New Productive Forces Need New Relations of Production</title>
    <link href="https://mochiaochen.github.io/en/writing/2026/07/new-productive-forces-need-new-relations-of-production/" rel="alternate" type="text/html"/>
    <published>2026-07-24T10:00:00+08:00</published>
    <updated>2026-07-24T10:00:00+08:00</updated>
    <id>https://mochiaochen.github.io/en/writing/2026/07/new-productive-forces-new-relations-of-production-en</id>
    <content type="html" xml:base="https://mochiaochen.github.io/en/writing/2026/07/new-productive-forces-need-new-relations-of-production/">&lt;p&gt;&lt;mark&gt;Every leap in productive forces ultimately requires new relations of production to sustain it.&lt;/mark&gt; Otherwise, new capabilities are trapped inside old structures and never fully released.&lt;/p&gt;

&lt;p&gt;Why did SpaceX succeed? It was not because NASA lacked the technology—it did not. The deeper reason is that NASA is a product of old relations of production. Its bureaucracy, contractor networks, and congressional budget cycles were designed for a previous generation of productive forces. Musk introduced a different set of relations—first-principles development, vertical integration, rapid iteration, and the unification of research and engineering—and used the same physics to produce a radically different result.&lt;/p&gt;

&lt;p&gt;This raises the real question: &lt;strong&gt;as a new productive force, what does agentic AI fundamentally change?&lt;/strong&gt;&lt;/p&gt;

&lt;h2 id=&quot;1-the-basic-unit-of-labour-has-changed&quot;&gt;1. The basic unit of labour has changed&lt;/h2&gt;

&lt;p&gt;In the agricultural age, the basic unit of production was &lt;strong&gt;people + land&lt;/strong&gt;. In the industrial age, it was &lt;strong&gt;people + machines&lt;/strong&gt;. In the digital age, it became &lt;strong&gt;engineers + code + data&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Today, as silicon intelligence learns to plan autonomously, execute independently, and produce continuously, that unit is becoming:&lt;/p&gt;

&lt;blockquote&gt;
  &lt;p&gt;&lt;strong&gt;Human judgement + AI agent systems&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;For the first time in history, cognitive labour itself can be automated at the level of execution. Earlier industrial revolutions replaced physical labour. This one is beginning to replace the execution layer of cognitive work: the execution involved in writing code, producing reports, and conducting analysis is increasingly being taken on by silicon.&lt;/p&gt;

&lt;p&gt;Human value will therefore become concentrated in two capabilities:&lt;/p&gt;

&lt;ol&gt;
  &lt;li&gt;&lt;strong&gt;the ability to ask the right questions;&lt;/strong&gt;&lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;the taste to judge what is valuable.&lt;/strong&gt;&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;mark&gt;When execution becomes cheap, questions and taste become expensive.&lt;/mark&gt;&lt;/p&gt;

&lt;h2 id=&quot;2-organisations-are-moving-towards-a-barbell-structure&quot;&gt;2. Organisations are moving towards a “barbell structure”&lt;/h2&gt;

&lt;p&gt;The optimal size of an organisation is shifting dramatically in two directions at once.&lt;/p&gt;

&lt;p&gt;At the research frontier, scale must keep growing. Training a new generation of models may require tens of billions of dollars in compute. This is driving the emergence of the &lt;strong&gt;NeoLab&lt;/strong&gt; as a new organisational form: it must carry scientific, engineering, market, and capital risk simultaneously. It is &lt;strong&gt;small and elite, yet extraordinarily capital-intensive&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;At the application and value-creation layer, meanwhile, the optimal scale is shrinking rapidly. &lt;strong&gt;FDE&lt;/strong&gt; is one signal; the &lt;strong&gt;OPC—One Person Company&lt;/strong&gt;—is the more extreme form. A product that once required a ten-person team can now be delivered by one person with exceptional judgement working alongside an AI agent system. Major Shlomo, working alone for six months, achieved an acquisition valuation of $25 million.&lt;/p&gt;

&lt;p&gt;At the organisational level, the new relations of production therefore resemble a barbell:&lt;/p&gt;

&lt;ul&gt;
  &lt;li&gt;at the top, a small number of enormous &lt;strong&gt;NeoLabs&lt;/strong&gt; concentrate research density, compute, and elite talent;&lt;/li&gt;
  &lt;li&gt;at the bottom, large numbers of small &lt;strong&gt;OPCs and FDEs&lt;/strong&gt; connect frontier capabilities flexibly to concrete industry settings;&lt;/li&gt;
  &lt;li&gt;in the middle sit traditional companies—especially those whose existence depends on headcount—and they will face the greatest pressure.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;mark&gt;Scale itself may not be the scarce resource of the future. High-density research and high-density judgement will be; low-density organisations caught between them are the most exposed.&lt;/mark&gt;&lt;/p&gt;

&lt;h2 id=&quot;3-capital-and-research-are-moving-from-financing-to-co-production&quot;&gt;3. Capital and research are moving from “financing” to “co-production”&lt;/h2&gt;

&lt;p&gt;The traditional venture-capital model says: &lt;strong&gt;I give you money; you give me a return within seven years.&lt;/strong&gt; It assumes that entrepreneurship is a process of rapid market validation, with short cycles and diversifiable risk.&lt;/p&gt;

&lt;p&gt;Research-driven entrepreneurship carries capital risk on an entirely different scale. The compute cost of training a frontier model, the time required to bring a brain–computer interface into clinical validation, and the path from a quantum-computing laboratory to industrial deployment all extend far beyond the limits of the traditional VC cycle.&lt;/p&gt;

&lt;p&gt;The new relations of production therefore require &lt;strong&gt;patient capital&lt;/strong&gt;: sovereign wealth funds, university endowments, and other long-term investors willing to allocate resources over ten or even twenty years. When this capital enters the production relationship, it is not merely financing a company. It is sharing with research entrepreneurs the risks of the production process itself.&lt;/p&gt;

&lt;p&gt;The relationships between Anthropic and Amazon, and between OpenAI and Microsoft, are no longer conventional investor–investee relationships. &lt;strong&gt;They are closer to structures of co-production.&lt;/strong&gt;&lt;/p&gt;

&lt;h2 id=&quot;4-compute-infrastructure-is-becoming-national&quot;&gt;4. Compute infrastructure is becoming national&lt;/h2&gt;

&lt;p&gt;The relations of production in the internet age were, to a large extent, transnational and detached from geography. A Silicon Valley company could serve users worldwide: the code lived in the cloud, and data appeared to have no borders.&lt;/p&gt;

&lt;p&gt;Compute is different. &lt;strong&gt;Compute is physical.&lt;/strong&gt; It depends on energy, chips fabricated with specialised processes, and geography. The state is therefore returning as a critical actor in the relations of production.&lt;/p&gt;

&lt;p&gt;America’s Stargate and China’s training grounds are, at their core, ways for states to organise the relations of production around compute infrastructure directly. Not since the railways and electrical grids of the industrial age has the state intervened in productive infrastructure with comparable force.&lt;/p&gt;

&lt;p&gt;The new relations of production thus gain another dimension. They concern not only firms, capital, and labour within markets, but also &lt;strong&gt;competition among major powers for control over the core factors of production&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;&lt;mark&gt;Once compute becomes infrastructure, technological competition also becomes a contest over energy, manufacturing, and state capacity.&lt;/mark&gt;&lt;/p&gt;

&lt;h2 id=&quot;5-the-distribution-of-value-is-undergoing-its-greatest-divergence&quot;&gt;5. The distribution of value is undergoing its greatest divergence&lt;/h2&gt;

&lt;p&gt;The old logic of value distribution was at least partly dispersed: firms created value, employees shared in it, and market competition spread some of it to users.&lt;/p&gt;

&lt;p&gt;The new productive structure creates an extreme duality of &lt;strong&gt;concentration and diffusion&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;At the productive-force layer—compute, models, and data—value becomes intensely concentrated. A small number of organisations controlling compute and frontier models enjoy near-monopolistic advantages.&lt;/p&gt;

&lt;p&gt;At the application layer, however, AI’s accessibility gives capabilities to people who could not previously enter certain fields. FDE can enable a single engineer to work deeply in medicine, law, or materials science and accomplish what once required an elite team of specialists.&lt;/p&gt;

&lt;p&gt;The result is a paradox:&lt;/p&gt;

&lt;blockquote&gt;
  &lt;p&gt;&lt;strong&gt;The concentration of the means of production is occurring alongside the diffusion of productive capacity.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Every major revolution in productive forces has displayed this pattern. The steam engine made a small number of people extraordinarily wealthy while bringing factories to every city. The same tension is returning today.&lt;/p&gt;

&lt;h2 id=&quot;6-institutional-choices-will-determine-the-direction-of-the-split&quot;&gt;6. Institutional choices will determine the direction of the split&lt;/h2&gt;

&lt;p&gt;Where this divergence ultimately leads depends on whether we make the right institutional choices:&lt;/p&gt;

&lt;ul&gt;
  &lt;li&gt;Can patient capital support genuine research?&lt;/li&gt;
  &lt;li&gt;Can broadly accessible infrastructure be built?&lt;/li&gt;
  &lt;li&gt;Can FDE capabilities be made available to more people?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;These questions are already larger than any individual or single organisation can answer alone. Yet they are becoming more important by the day.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Agentic AI is not merely another upgrade to our tools.&lt;/strong&gt; It is rewriting the relationships among labour, organisations, capital, the state, and distribution. The new productive force is already here. The unresolved question is whether we can create new relations of production quickly enough to sustain it.&lt;/p&gt;

&lt;p&gt;&lt;mark&gt;If we cannot, new capabilities will only reinforce old structures. If we can, they may finally become the productive forces of a new era.&lt;/mark&gt;&lt;/p&gt;
</content>
  </entry>
  
  <entry>
    <title>上香时代：一代人的精神退场</title>
    <link href="https://mochiaochen.github.io/writing/2026/07/the-age-of-incense/" rel="alternate" type="text/html"/>
    <published>2026-07-22T17:30:00+08:00</published>
    <updated>2026-07-22T17:30:00+08:00</updated>
    <id>https://mochiaochen.github.io/writing/2026/07/the-age-of-incense</id>
    <content type="html" xml:base="https://mochiaochen.github.io/writing/2026/07/the-age-of-incense/">&lt;h2 id=&quot;在上班与上进之间&quot;&gt;在上班与上进之间&lt;/h2&gt;

&lt;p&gt;2023 年春天，一条数据在社交平台上流转：携程与美团的寺庙景区订票量同比暴涨超过三倍，其中近半数购票者是 90 后与 00 后。北京雍和宫的头柱香需要凌晨排队，杭州灵隐寺的十八籽手串一度卖到断货，五台山、普陀山、峨眉山的旺季提前了整整一个月。与此同时，各大电商平台的塔罗牌销量在两年内翻了四倍，八字排盘类小程序的月活用户以千万计，MBTI 人格测试从一个心理学缩写变成了社交货币（自我介绍的开场白从「我是做什么的」变成了「我是 INFP」），而一款名为「电子木鱼」的手机应用以其荒诞的简洁征服了整整一代人：点击屏幕，木鱼响一声，功德加一。&lt;/p&gt;

&lt;p&gt;一个句子像种子一样从互联网的腐殖质里长出来，迅速成为这个时代最精准的自嘲：&lt;/p&gt;

&lt;blockquote&gt;
  &lt;p&gt;「在上班和上进之间，我选择了上香。」&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;&lt;strong&gt;笑声的底层是一份精确的成本收益核算。&lt;/strong&gt;上班的回报在缩水（实际工资增速放缓，35 岁危机悬顶，裁员的消息以日为周期更新），上进的通道在变窄（考研报名人数从 2019 年的 290 万涨到 2024 年的 438 万，录取率持续走低；国考报名人数逼近四百万，近百人争一个岗位），而上香的成本几乎为零，回报虽然是虚拟的，但至少没有人会在香炉前收到一封写着「感谢您的参与，您未通过本轮筛选」的邮件。当正统的晋升通道拥堵到接近瘫痪，人总要找一个地方安放自己对未来的期待，哪怕那个地方供奉的是一尊泥塑。&lt;/p&gt;

&lt;p&gt;涂尔干在 19 世纪末观察法国社会时，为这类现象锻造了一个至今仍然锋利的概念：失范（anomie）。在《自杀论》中，他写道：&lt;/p&gt;

&lt;blockquote&gt;
  &lt;p&gt;“Anomie, therefore, is a regular and specific factor in suicide in our modern societies; one of the springs from which the annual contingent feeds.”&lt;sup id=&quot;fnref:1&quot; role=&quot;doc-noteref&quot;&gt;&lt;a href=&quot;#fn:1&quot; class=&quot;footnote&quot; rel=&quot;footnote&quot;&gt;1&lt;/a&gt;&lt;/sup&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;失范的准确含义并非「没有规范」，而是「旧规范已经瓦解，新规范尚未建立」的过渡状态。社会为个体提供的意义框架失效了，而替代框架还没有长出来，于是个体悬浮在真空之中，既不知道自己的努力指向何方，也不知道用什么标准来衡量自己的成败。涂尔干诊断的是 19 世纪欧洲从传统社会向工业社会的急速转型，但他所刻画的那种「期望与现实之间的制度性裂缝」，放在今天的中国，几乎不需要做任何翻译。&lt;/p&gt;

&lt;p&gt;寺庙热、塔罗、八字、MBTI、电子木鱼，表面上五花八门，底层结构是同一个：它们都在为一种丧失了制度支撑的个人未来感提供廉价的叙事替代品。MBTI 告诉你「你是谁」（在一个岗位不再能定义身份的年代），塔罗告诉你「接下来会怎样」（在一个规划已经失去可信度的年代），而上香则把全部的不确定性交给了一个不会拒绝你的对象。&lt;mark&gt;它们共同填充的那个空洞，有一个更古老的名字。&lt;/mark&gt;&lt;/p&gt;

&lt;h2 id=&quot;一千七百年前的精神退出&quot;&gt;一千七百年前的精神退出&lt;/h2&gt;

&lt;p&gt;公元 3 世纪，中国历史上曾经发生过一次惊人相似的精神地震。&lt;/p&gt;

&lt;p&gt;东汉末年，经学与察举构成的晋升锦标赛运转了四百年之后，终于走向全面崩解。经学的权威在党锢之祸中被政治摧毁，察举的公正性在门阀垄断中被利益腐蚀，儒家的道德话语（「名教」）与权力实践之间的裂缝大到再也无法修辞性地弥合。司马氏以禅让之名行篡夺之实，把「忠孝仁义」的最后一层遮羞布扯了下来。当正统通道的回报塌陷，当官方话语的信用破产，精英的力气从外转向内，从经世转向玄谈，从入仕转向山林。&lt;/p&gt;

&lt;p&gt;余英时在《士与中国文化》中对这一转向做过精密的思想史分析。他指出，魏晋玄学的兴起，本质上是名教（制度化的儒家伦理）失去权威之后，士人阶层寻找替代性精神安顿的一场集体实验。他写道：&lt;/p&gt;

&lt;blockquote&gt;
  &lt;p&gt;「魏晋士大夫对名教的态度，大体上可以分成三派：一是维护名教，二是在名教中注入新的精神，三是完全反名教而归于自然。」&lt;sup id=&quot;fnref:2&quot; role=&quot;doc-noteref&quot;&gt;&lt;a href=&quot;#fn:2&quot; class=&quot;footnote&quot; rel=&quot;footnote&quot;&gt;2&lt;/a&gt;&lt;/sup&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;这三条路径的分布，像一张光谱一样覆盖了整个精英阶层。维护名教的是那些仍然在体制内运转的士族官僚，他们的对应物是今天仍然全力以赴冲击考公考编的年轻人。「在名教中注入新精神」的是王弼、郭象这样试图调和儒道的思想家，他们的对应物或许是今天在体制边缘寻找「意义创业」的知识青年。而「完全反名教归于自然」的那一派，其最激烈的代言人，名叫嵇康。&lt;/p&gt;

&lt;p&gt;嵇康留下了一封绝交信，中国文学史上最著名的绝交信。公元 261 年前后，他的老友山涛（字巨源）被任命为吏部选曹郎，推荐嵇康出仕。嵇康以这封《与山巨源绝交书》回绝，信中罗列了自己不适合做官的种种理由，语气在戏谑与决绝之间游走，到最后突然亮出了一把匕首般的短句：&lt;/p&gt;

&lt;blockquote&gt;
  &lt;p&gt;「每非汤武而薄周孔，在人间不止此事，会显世教所不容。」&lt;sup id=&quot;fnref:3&quot; role=&quot;doc-noteref&quot;&gt;&lt;a href=&quot;#fn:3&quot; class=&quot;footnote&quot; rel=&quot;footnote&quot;&gt;3&lt;/a&gt;&lt;/sup&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;非议汤武，鄙薄周孔。这八个字在当时的杀伤力，大约相当于今天一个人在公开场合说出「我认为考公是对生命的浪费」。它否定的不是某一项具体政策，它否定的是整个意义体系的合法性：你们用来衡量成功的那套标准，我拒绝承认。嵇康自知这番话「会显世教所不容」，他是对的。两年后，司马昭将他处死。刑场上嵇康索琴弹了一曲《广陵散》，然后从容赴死。&lt;/p&gt;

&lt;p&gt;这个故事里最值得注意的细节，恰恰是嵇康的选择方式。他没有起兵、没有上书、没有组织任何形式的反对运动。他的反抗纯粹是精神性的：拒绝出仕，拒绝配合，拒绝用体制的语言来描述自己的生活。&lt;strong&gt;他选择了退出。&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;2020 年代的中国青年做了几乎完全相同的事。他们的「非汤武而薄周孔」是不转发、不点赞、不表态的沉默，是全职儿女、慢就业、二战三战的拖延，是「45 度人生」的刻意减速。他们没有组织、没有领袖、没有纲领，他们只是静静地、一个一个地，从赛道的边缘走开了。如果说嵇康的退出是一个人的决绝，那么当数以百万计的年轻人同时做出这个选择的时候，它就不再是个人气质的表达，它就成了一组宏观数据。而当这种退出延伸到最私密的领域，延伸到「不结婚、不生育」的选择时，它就成了一场无声的全民公决。&lt;/p&gt;

&lt;h2 id=&quot;魏晋镜像的边界&quot;&gt;魏晋镜像的边界&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;类比照明，但类比需要纪律。&lt;/strong&gt;魏晋的镜子能照出 2020 年代中国的哪些面向，又在哪里折射出误导性的光？&lt;/p&gt;

&lt;p&gt;相似之处是深层结构性的。两个时代都处于一套晋升锦标赛的末期疲劳阶段：汉代的经学察举制与当代的教育军备竞赛加公务员考试，都是以极高的个人投入换取有限的上升名额，都在运行数代人之后遭遇边际回报的急剧递减。两个时代的官方话语都出现了严重的信用赤字：名教的虚伪与今天「正能量」话语和个人体感之间的裂缝，都迫使精英阶层在公开表达与私下判断之间维持一道日益昂贵的防火墙。两个时代的精英都把力气从公共参与收回到了个体的审美与内在性：竹林七贤的饮酒、弹琴、清谈，与今天的寺庙游、塔罗牌、播客听书，在功能上高度同构。&lt;/p&gt;

&lt;p&gt;刘义庆在《世说新语》里记载了一个细节，几乎可以直接贴进今天的社交媒体。王子猷（王徽之）居住在山阴，一夜大雪，他半夜醒来，开窗饮酒，忽然想起朋友戴安道住在剡溪，于是连夜乘舟前往。走了一整夜，到了戴家门口，却不进去，转身就走。人问其故，他说：&lt;/p&gt;

&lt;blockquote&gt;
  &lt;p&gt;「吾本乘兴而行，兴尽而返，何必见戴？」&lt;sup id=&quot;fnref:4&quot; role=&quot;doc-noteref&quot;&gt;&lt;a href=&quot;#fn:4&quot; class=&quot;footnote&quot; rel=&quot;footnote&quot;&gt;4&lt;/a&gt;&lt;/sup&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;「乘兴而行，兴尽而返」，这八个字是整个魏晋美学的微缩：过程即目的，体验即意义，不为结果服务的行动才是自由的行动。今天的年轻人管这种状态叫「听从内心」或者「做自己」，他们在特种兵式旅游、City Walk、搭子社交中实践的，正是某种降维版的「乘兴而行」。区别在于，王子猷能够「乘兴」，是因为他有琅琊王氏的庄园可以依托；今天的年轻人的「乘兴」，往往在月底的信用卡账单面前戛然而止。&lt;/p&gt;

&lt;p&gt;差异之处同样重要，忽视差异的类比必然滑向抒情。&lt;/p&gt;

&lt;p&gt;第一个差异在于退路。魏晋的世族拥有自给自足的庄园经济，「归隐山林」在物质层面是一个真实的选项，陶渊明可以「采菊东篱下」，前提是他有那道东篱。今天的青年没有庄园，他们有的是十几平米的出租屋和三十分钟送达的外卖。现代社会的分工体系不允许任何人真正退出：你可以精神上躺平，但你的房租、社保、手机话费每个月都在要求你以某种方式接入系统。&lt;mark&gt;退出是姿态性的，接入是结构性的。&lt;/mark&gt;这意味着当代的「躺平」比魏晋的「归隐」在心理上更为折磨，因为它是一种不彻底的退出，一种在参与和拒绝之间永久摆荡的状态。&lt;/p&gt;

&lt;p&gt;第二个差异在于后果。魏晋清谈之后，是三百年的分裂与战乱（五胡十六国、南北朝），其间伴随着巨大的人口损失和文明的局部断裂。这是许多使用魏晋类比的论者暗中指向的恐怖结局：精神的失范预示着政治的崩解。但这个推论在方法论上站不住脚。魏晋的政治崩溃有其独立的、充分的结构性原因（军事贵族的崛起、中央军事力量的真空、北方民族的南迁压力），清谈与崩溃之间的因果关系从未被严格建立。&lt;strong&gt;精神的失范可以是崩溃的前兆，也可以仅仅是社会在承压下的自我调适。&lt;/strong&gt;从它到政治崩解之间，隔着无数个需要独立验证的中间环节。&lt;/p&gt;

&lt;p&gt;那么，今天的精神失范更可能通向哪里？&lt;/p&gt;

&lt;h2 id=&quot;更接近日本的未来&quot;&gt;更接近日本的未来&lt;/h2&gt;

&lt;p&gt;一个更晚近、也更具可比性的参照系，是 1990 年代之后的日本。&lt;/p&gt;

&lt;p&gt;大前研一在 2015 年出版的《低欲望社会》中，对泡沫破裂后日本年轻一代的精神状态做了一番近乎人类学的观察。他写道：&lt;/p&gt;

&lt;blockquote&gt;
  &lt;p&gt;「日本年轻人没有欲望、没有梦想、没有干劲。……不想出人头地，不想拥有汽车和奢侈品，消费意愿降至最低。」&lt;sup id=&quot;fnref:5&quot; role=&quot;doc-noteref&quot;&gt;&lt;a href=&quot;#fn:5&quot; class=&quot;footnote&quot; rel=&quot;footnote&quot;&gt;5&lt;/a&gt;&lt;/sup&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;大前研一把这种状态命名为「低欲望社会」，它的特征包括：年轻人拒绝承担房贷、推迟甚至放弃结婚与生育、消费降级为日常习惯、对晋升丧失兴趣、转而追求「小确幸」式的私人满足。这些特征的清单拿过来描述 2024 年的中国城市青年，几乎不需要改动一个字。&lt;/p&gt;

&lt;p&gt;日本的经验给出了一个关键的实证反例：精神失范并不必然导向政治崩解。泡沫破裂后的日本经历了「失去的三十年」，GDP 增速徘徊于零线附近，人口持续萎缩，但社会秩序始终稳定，民主体制照常运转，犯罪率甚至逐年下降。一个低欲望的社会完全可以是一个安静、有序、缓慢老去的社会。日本证明了现代性的一种可能：不必以激烈的断裂收场，而以漫长的停滞和温和的收缩来消化一代人的失望。&lt;/p&gt;

&lt;p&gt;用这个框架来看，2020 年代中国的精神图景更可能的走向，与其说是魏晋式的王朝解体，不如说是日本式的低欲望社会。内卷、躺平、上香、电子木鱼，这些现象的共同特征是退缩而非对抗，是降低期待而非掀翻赌桌。它们指向的是现代性的到达，而非文明的终结。当一个社会足够富裕、足够安全、足够原子化，以至于个体可以在不依赖任何集体行动的前提下独自完成对主流价值的拒绝时，这种拒绝就天然地趋向安静而非暴烈。&lt;/p&gt;

&lt;p&gt;但日本类比同样需要纪律。中日之间横着三道裂缝。第一，日本变老时人均 GDP 已逾四万美元，社会安全网基本完备，医疗与养老覆盖全民；中国在人均 1.3 万美元处变老，近三亿农民工没有城镇职工养老保险，安全网的缺口意味着「低欲望」缺乏日本式的物质底板。第二，日本的停滞发生在美国安全伞之内，贸易渠道大体畅通；中国的停滞将在关税墙与技术封锁之中进行，外部环境截然不同。第三，也是最根本的一条：日本有选举制度作为泄压阀，三十年间首相更替十数次，政策路线随民意摆动，失望可以通过投票表达，哪怕表达的效果有限。中国没有这层缓冲。当失望既无法通过退出（社会结构不允许真正退出）也无法通过呼喊（政治结构不提供呼喊的渠道）来释放时，它就只能沉淀。沉淀的失望不会消失，它只是变得沉默而稠密，等待某个不可预测的时刻，以不可预测的方式，寻找出口。&lt;/p&gt;

&lt;p&gt;赫希曼（Albert Hirschman）在 1970 年提出的经典框架在此处具有近乎残酷的解释力。面对令人不满的组织或体制，个体有三种选择：退出（exit）、呼喊（voice）、忠诚（loyalty）。赫希曼写道：&lt;/p&gt;

&lt;blockquote&gt;
  &lt;p&gt;“The exit option is widely held to be uniquely important in keeping sellers alert. It is the threat of exit that acts as the primary restraining force.”&lt;sup id=&quot;fnref:6&quot; role=&quot;doc-noteref&quot;&gt;&lt;a href=&quot;#fn:6&quot; class=&quot;footnote&quot; rel=&quot;footnote&quot;&gt;6&lt;/a&gt;&lt;/sup&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;退出的威胁是迫使组织自我纠正的首要力量。但在一个地理退出成本极高（润的门槛）、政治呼喊渠道极窄、而忠诚已经被透支的社会里，三条路都半堵着。于是出现了第四种选项，赫希曼的模型里没有命名的那一种：留在原地，但把音量调到最低。不走，不喊，不效忠，只是安静地降低自己对一切的期待。这就是今天正在发生的事。&lt;mark&gt;它比退出更沉默，比呼喊更安全，比忠诚更诚实。&lt;/mark&gt;&lt;/p&gt;

&lt;h2 id=&quot;附近的消失&quot;&gt;「附近」的消失&lt;/h2&gt;

&lt;p&gt;在所有关于当代中国精神状态的分析中，人类学家项飙提出的一个概念最具穿透力。他称之为「附近的消失」。&lt;/p&gt;

&lt;p&gt;项飙在与许知远的对谈中阐述这个概念时指出，当代中国人的生活在空间上呈现出一种奇特的哑铃形结构：一端是极度私密的自我（手机屏幕、内心世界），另一端是极度宏大的集体想象（国家、民族、全球格局），而中间那个层次，即「附近」，也就是你所居住的社区、你的邻居、你日常行走的街道、你与之发生面对面互动的那些具体的人和场所，正在系统性地消失。他这样概括：&lt;/p&gt;

&lt;blockquote&gt;
  &lt;p&gt;「我们现在的注意力分配出现了一个很大的问题。对自我的关注和对世界的关注都很多，但是中间的’附近’这个层面，作为一个生活的基本感知单元，被掏空了。」&lt;sup id=&quot;fnref:7&quot; role=&quot;doc-noteref&quot;&gt;&lt;a href=&quot;#fn:7&quot; class=&quot;footnote&quot; rel=&quot;footnote&quot;&gt;7&lt;/a&gt;&lt;/sup&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;&lt;strong&gt;「附近」的消失不是一个空间问题，它是一个意义问题。&lt;/strong&gt;当你不认识你的邻居，不关心你所在街道的变化，不参与任何社区层面的集体生活时，你与世界之间就缺少了一个最基本的中介层。宏大叙事（「民族复兴」「大国崛起」）无论多么激昂，它都无法直接兑付为个体的日常意义感；而纯粹的私人生活（消费、娱乐、自我关照）无论多么精致，它都缺乏来自他者的确认与回响。&lt;mark&gt;人需要「附近」，因为「附近」是意义生长的土壤：你在邻里中被认识、被需要、被记住，你的行动在一个可触及的范围内产生可见的后果。&lt;/mark&gt;&lt;/p&gt;

&lt;p&gt;项飙的诊断与涂尔干的失范概念在深层互相呼应。涂尔干认为，现代社会的失范根源在于传统共同体的瓦解；项飙则精确地指出了瓦解发生的具体层次：非最亲密的家庭（它仍然存在），亦非最遥远的国家（它无处不在），乃是中间那个让个体与社会产生有机联系的「附近」。在中国的语境下，这种消失有其特殊的加速器：城市化把数亿人从熟人社会连根拔起，种进陌生人构成的小区；平台经济把一切日常需求（吃饭、出行、购物、社交）从街道搬到屏幕上，街道失去了作为公共生活场所的功能；而社区层面的自治组织在体制设计中长期缺位，居委会更多是行政末梢而非居民自治的载体。于是「附近」像一块被上下两端同时抽走的布，越拉越薄，终至透明。&lt;/p&gt;

&lt;p&gt;悬浮由此成为这一代人的存在底色。你可以在手机上为国产大飞机的首飞热泪盈眶，转身却叫不出对面邻居的名字。你可以对中美芯片博弈了如指掌，却不知道自己所住的楼层水管归谁修。宏大叙事提供自豪感，私人世界提供安慰，但两者之间没有桥梁。这种悬浮状态的心理代价是一种弥散性的不真实感：一切都在运转，你也在其中运转，但你无法确定自己与这台机器之间的关系到底是什么。你既不是它的零件（你随时可以被替换），也不是它的主人（你没有任何决定权），你只是刚好在场。&lt;/p&gt;

&lt;h2 id=&quot;沉默的拒绝&quot;&gt;沉默的拒绝&lt;/h2&gt;

&lt;p&gt;新清谈的表层是消费现象：寺庙门票、塔罗牌销量、MBTI 标签、电子木鱼的下载量。中层是社会心理现象：正统通道的回报塌陷之后，一代人把力气从外部竞争收回到内在体验，从长期规划收回到即时感受。深层是结构性的意义危机：旧的社会契约（努力就能上升）失效之后，新的契约尚未建立，「附近」被掏空，个体悬浮在私人与国家之间的真空中，靠碎片化的精神消费来维持意义的最低供给。&lt;/p&gt;

&lt;p&gt;魏晋的类比照亮了这个局面的一半：当官方话语的信用破产、晋升锦标赛的边际回报归零时，精英的精神转向具有跨越一千七百年的结构相似性。但类比必须止步于此。魏晋之后的三百年分裂有其独立的军事与政治原因，不能从精神失范直接推导。更可能的参照系是日本：低欲望社会的缓慢老去，安静、有序，但也沉闷、停滞。然而中日之间的三道裂缝（收入水平、外部环境、泄压机制）意味着中国版的低欲望社会将会更粗粝、更紧张、更缺少缓冲。&lt;/p&gt;

&lt;p&gt;嵇康在赴死之前弹完最后一曲，说了一句话：「《广陵散》于今绝矣。」&lt;sup id=&quot;fnref:8&quot; role=&quot;doc-noteref&quot;&gt;&lt;a href=&quot;#fn:8&quot; class=&quot;footnote&quot; rel=&quot;footnote&quot;&gt;8&lt;/a&gt;&lt;/sup&gt;一首曲子的失传，在他看来比一条命的终结更值得惋惜。这种把审美置于生存之上的极端姿态，在今天当然不可复制。但他的那个核心动作，那个用沉默、用不配合、用转身离开来表达的拒绝，正在以亿为单位被复制着。不同的是，嵇康知道自己在拒绝什么。今天的许多年轻人，他们不一定能说清自己在拒绝什么，他们只是感到了倦。倦于赛道、倦于表演、倦于那种「你再努力一点就可以了」的许诺。&lt;/p&gt;

&lt;p&gt;于是他们上香。功德 +1。&lt;/p&gt;

&lt;hr /&gt;

&lt;p&gt;&lt;strong&gt;注释&lt;/strong&gt;&lt;/p&gt;

&lt;div class=&quot;footnotes&quot; role=&quot;doc-endnotes&quot;&gt;
  &lt;ol&gt;
    &lt;li id=&quot;fn:1&quot; role=&quot;doc-endnote&quot;&gt;
      &lt;p&gt;Émile Durkheim, &lt;em&gt;Suicide: A Study in Sociology&lt;/em&gt;, trans. John A. Spaulding &amp;amp; George Simpson, The Free Press, 1951 (orig. 1897), p. 258. &lt;a href=&quot;#fnref:1&quot; class=&quot;reversefootnote&quot; role=&quot;doc-backlink&quot;&gt;&amp;#8617;&lt;/a&gt;&lt;/p&gt;
    &lt;/li&gt;
    &lt;li id=&quot;fn:2&quot; role=&quot;doc-endnote&quot;&gt;
      &lt;p&gt;余英时，《士与中国文化》，上海人民出版社，2003 年，第 335 页。 &lt;a href=&quot;#fnref:2&quot; class=&quot;reversefootnote&quot; role=&quot;doc-backlink&quot;&gt;&amp;#8617;&lt;/a&gt;&lt;/p&gt;
    &lt;/li&gt;
    &lt;li id=&quot;fn:3&quot; role=&quot;doc-endnote&quot;&gt;
      &lt;p&gt;嵇康，《与山巨源绝交书》，参见《晋书·嵇康传》。 &lt;a href=&quot;#fnref:3&quot; class=&quot;reversefootnote&quot; role=&quot;doc-backlink&quot;&gt;&amp;#8617;&lt;/a&gt;&lt;/p&gt;
    &lt;/li&gt;
    &lt;li id=&quot;fn:4&quot; role=&quot;doc-endnote&quot;&gt;
      &lt;p&gt;刘义庆，《世说新语·任诞》。 &lt;a href=&quot;#fnref:4&quot; class=&quot;reversefootnote&quot; role=&quot;doc-backlink&quot;&gt;&amp;#8617;&lt;/a&gt;&lt;/p&gt;
    &lt;/li&gt;
    &lt;li id=&quot;fn:5&quot; role=&quot;doc-endnote&quot;&gt;
      &lt;p&gt;大前研一，《低欲望社会：「丧失大志时代」的新·国富论》（低欲望社会：「大志なき时代」の新·国富论），小学馆新书，2015 年。 &lt;a href=&quot;#fnref:5&quot; class=&quot;reversefootnote&quot; role=&quot;doc-backlink&quot;&gt;&amp;#8617;&lt;/a&gt;&lt;/p&gt;
    &lt;/li&gt;
    &lt;li id=&quot;fn:6&quot; role=&quot;doc-endnote&quot;&gt;
      &lt;p&gt;Albert O. Hirschman, &lt;em&gt;Exit, Voice, and Loyalty: Responses to Decline in Firms, Organizations, and States&lt;/em&gt;, Harvard University Press, 1970, p. 21. &lt;a href=&quot;#fnref:6&quot; class=&quot;reversefootnote&quot; role=&quot;doc-backlink&quot;&gt;&amp;#8617;&lt;/a&gt;&lt;/p&gt;
    &lt;/li&gt;
    &lt;li id=&quot;fn:7&quot; role=&quot;doc-endnote&quot;&gt;
      &lt;p&gt;项飙，与许知远对谈，「十三邀」第四季，2019 年。此处为大意转述。 &lt;a href=&quot;#fnref:7&quot; class=&quot;reversefootnote&quot; role=&quot;doc-backlink&quot;&gt;&amp;#8617;&lt;/a&gt;&lt;/p&gt;
    &lt;/li&gt;
    &lt;li id=&quot;fn:8&quot; role=&quot;doc-endnote&quot;&gt;
      &lt;p&gt;嵇康，临刑语，参见《世说新语·雅量》。 &lt;a href=&quot;#fnref:8&quot; class=&quot;reversefootnote&quot; role=&quot;doc-backlink&quot;&gt;&amp;#8617;&lt;/a&gt;&lt;/p&gt;
    &lt;/li&gt;
  &lt;/ol&gt;
&lt;/div&gt;
</content>
  </entry>
  
  <entry>
    <title>The Age of Incense: A Generation&apos;s Quiet Retreat</title>
    <link href="https://mochiaochen.github.io/en/writing/2026/07/the-age-of-incense/" rel="alternate" type="text/html"/>
    <published>2026-07-22T17:30:00+08:00</published>
    <updated>2026-07-22T17:30:00+08:00</updated>
    <id>https://mochiaochen.github.io/en/writing/2026/07/the-age-of-incense-en</id>
    <content type="html" xml:base="https://mochiaochen.github.io/en/writing/2026/07/the-age-of-incense/">&lt;h2 id=&quot;between-work-and-getting-ahead&quot;&gt;Between work and getting ahead&lt;/h2&gt;

&lt;p&gt;In the spring of 2023, a set of figures circulated on social media: bookings for temple attractions on Ctrip and Meituan had more than tripled year on year, and nearly half of the visitors were born in the 1990s or 2000s. Those seeking the first incense offering of the day at Beijing’s Yonghe Temple had to queue before dawn. Lingyin Temple in Hangzhou briefly sold out of its eighteen-seed bracelets. The peak season at Mount Wutai, Mount Putuo, and Mount Emei arrived a full month early. Meanwhile, tarot-card sales on major e-commerce platforms quadrupled in two years; mini-programs generating &lt;em&gt;bazi&lt;/em&gt; charts counted tens of millions of monthly active users; MBTI grew from an abbreviation in psychology into a form of social currency—the opening line of a self-introduction changed from “what I do” to “I am an INFP”; and a mobile app called “Digital Wooden Fish” conquered an entire generation with its absurd simplicity: tap the screen, hear the wooden fish, gain one point of merit.&lt;/p&gt;

&lt;p&gt;A sentence sprouted like a seed from the internet’s humus and quickly became the most precise self-mockery of the age:&lt;/p&gt;

&lt;blockquote&gt;
  &lt;p&gt;“Between working and working my way up, I chose to offer incense.”&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;&lt;strong&gt;Beneath the laughter lies a precise cost-benefit calculation.&lt;/strong&gt; The returns from working are shrinking: real wage growth is slowing, the crisis at thirty-five hangs overhead, and layoff news updates daily. The route to getting ahead is narrowing: postgraduate entrance-exam registrations rose from 2.9 million in 2019 to 4.38 million in 2024 as admission rates continued to fall; applications for the national civil-service examination approached four million, with nearly a hundred people competing for one post. Offering incense, by contrast, costs almost nothing. Its return may be imaginary, but at least no one receives an email at the censer saying, “Thank you for your participation. You have not passed this round.” When conventional routes of advancement are congested almost to paralysis, people need somewhere to place their hopes for the future—even if that place houses a clay idol.&lt;/p&gt;

&lt;p&gt;Observing French society at the end of the nineteenth century, Émile Durkheim forged a concept for phenomena of this kind that remains sharp today: &lt;strong&gt;anomie&lt;/strong&gt;. In &lt;em&gt;Suicide&lt;/em&gt;, he wrote:&lt;/p&gt;

&lt;blockquote&gt;
  &lt;p&gt;“Anomie, therefore, is a regular and specific factor in suicide in our modern societies; one of the springs from which the annual contingent feeds.”&lt;sup id=&quot;fnref:1-en&quot; role=&quot;doc-noteref&quot;&gt;&lt;a href=&quot;#fn:1-en&quot; class=&quot;footnote&quot; rel=&quot;footnote&quot;&gt;1&lt;/a&gt;&lt;/sup&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Anomie does not precisely mean “the absence of norms.” It is a transitional condition in which old norms have broken down and new ones have yet to form. The framework of meaning that society offers the individual has failed, while no replacement has grown in its place. The individual is suspended in a vacuum, knowing neither where effort is meant to lead nor by what standard success and failure should be measured. Durkheim diagnosed Europe’s rapid transition from traditional to industrial society in the nineteenth century, but his “institutional gap between expectations and reality” needs almost no translation in contemporary China.&lt;/p&gt;

&lt;p&gt;Temple fever, tarot, &lt;em&gt;bazi&lt;/em&gt;, MBTI, and the digital wooden fish look wildly different on the surface. Their underlying structure is the same: each provides a cheap narrative substitute for a sense of personal future that has lost its institutional support. MBTI tells you “who you are” in an age when a job no longer defines identity. Tarot tells you “what happens next” in an age when planning has lost credibility. Offering incense hands all uncertainty to an object that will never reject you. &lt;mark&gt;The void they fill has a much older name.&lt;/mark&gt;&lt;/p&gt;

&lt;h2 id=&quot;a-spiritual-withdrawal-seventeen-centuries-ago&quot;&gt;A spiritual withdrawal seventeen centuries ago&lt;/h2&gt;

&lt;p&gt;In the third century, Chinese history experienced a strikingly similar spiritual earthquake.&lt;/p&gt;

&lt;p&gt;After four hundred years, the tournament of advancement built from Confucian classics and the recommendation system of the Eastern Han finally collapsed. The authority of classical learning was politically destroyed in the Partisan Prohibitions; the fairness of recommendation was corroded by the monopoly of great clans; and the gulf between Confucian moral discourse—&lt;em&gt;mingjiao&lt;/em&gt;, the orthodox code—and the exercise of power grew too wide for rhetoric to conceal. The Sima clan usurped the throne in the name of abdication and tore away the final veil over “loyalty, filial piety, benevolence, and righteousness.” As returns from the orthodox route collapsed and official discourse lost credibility, the elite turned its energy inward: from statecraft to metaphysical conversation, and from official service to the mountains.&lt;/p&gt;

&lt;p&gt;In &lt;em&gt;The Scholar and Chinese Culture&lt;/em&gt;, Yu Yingshi offered a precise intellectual history of this turn. The rise of Wei-Jin &lt;em&gt;xuanxue&lt;/em&gt;, he argued, was essentially a collective experiment in which the scholar-gentry sought alternative spiritual settlement after institutionalised Confucian ethics lost authority. He wrote:&lt;/p&gt;

&lt;blockquote&gt;
  &lt;p&gt;“The attitudes of Wei-Jin scholar-officials toward &lt;em&gt;mingjiao&lt;/em&gt; can broadly be divided into three schools: those who defended it, those who infused it with a new spirit, and those who rejected it entirely and returned to nature.”&lt;sup id=&quot;fnref:2-en&quot; role=&quot;doc-noteref&quot;&gt;&lt;a href=&quot;#fn:2-en&quot; class=&quot;footnote&quot; rel=&quot;footnote&quot;&gt;2&lt;/a&gt;&lt;/sup&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;These three paths spread across the elite like a spectrum. Those defending &lt;em&gt;mingjiao&lt;/em&gt; were the aristocratic officials still operating within the system; their counterparts today are young people still giving everything to civil-service and public-institution examinations. Those “infusing &lt;em&gt;mingjiao&lt;/em&gt; with a new spirit” were thinkers such as Wang Bi and Guo Xiang, who sought a reconciliation between Confucianism and Daoism; their counterparts might be young intellectuals pursuing “meaning entrepreneurship” at the system’s edge. The most uncompromising spokesman for the third path—complete rejection of &lt;em&gt;mingjiao&lt;/em&gt; and return to nature—was Ji Kang.&lt;/p&gt;

&lt;p&gt;Ji Kang left behind the most famous letter of severance in Chinese literary history. Around 261, his old friend Shan Tao, courtesy name Juyuan, was appointed to a personnel office and recommended Ji Kang for government service. Ji refused in his &lt;em&gt;Letter Breaking Off Relations with Shan Juyuan&lt;/em&gt;. He listed his many reasons for being unsuited to office, moving between mockery and resolve, before suddenly drawing a dagger of a sentence:&lt;/p&gt;

&lt;blockquote&gt;
  &lt;p&gt;“I have always condemned Tang and Wu and held the Duke of Zhou and Confucius in contempt. This is not the only such matter in my life; once revealed, it cannot be tolerated by the teachings of the age.”&lt;sup id=&quot;fnref:3-en&quot; role=&quot;doc-noteref&quot;&gt;&lt;a href=&quot;#fn:3-en&quot; class=&quot;footnote&quot; rel=&quot;footnote&quot;&gt;3&lt;/a&gt;&lt;/sup&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;To condemn Tang and Wu and disdain Zhou and Confucius had roughly the force, in that time, of someone announcing publicly today: “I believe taking the civil-service examination is a waste of life.” It rejected no particular policy. It rejected the legitimacy of an entire system of meaning: I refuse to recognise the standards by which you measure success. Ji knew that these words would be “intolerable to the teachings of the age.” He was right. Two years later, Sima Zhao had him executed. At the place of execution, Ji played &lt;em&gt;Guangling San&lt;/em&gt; on the &lt;em&gt;qin&lt;/em&gt; and calmly went to his death.&lt;/p&gt;

&lt;p&gt;The most revealing detail is the form of his choice. Ji raised no army, submitted no memorial, and organised no opposition. His resistance was wholly spiritual: he refused office, refused cooperation, and refused to describe his life in the language of the system. &lt;strong&gt;He chose to withdraw.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Young Chinese people in the 2020s have done almost exactly the same. Their “condemning Tang and Wu and scorning Zhou and Confucius” is the silence of not reposting, liking, or taking a position; the postponement of becoming full-time children, slow employment, and repeated examination attempts; the deliberate deceleration of a “45-degree life.” They have no organisation, leader, or programme. Quietly, one by one, they simply walk away from the edge of the track. Ji Kang’s withdrawal was one person’s resolve. When millions make the same choice, it is no longer an expression of individual temperament but a set of macroeconomic data. And when that withdrawal reaches into the most intimate domains—the choice not to marry or have children—it becomes a silent nationwide referendum.&lt;/p&gt;

&lt;h2 id=&quot;the-limits-of-the-wei-jin-mirror&quot;&gt;The limits of the Wei-Jin mirror&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Analogies illuminate, but analogies require discipline.&lt;/strong&gt; Which aspects of China in the 2020s can the Wei-Jin mirror reveal, and where does it refract misleading light?&lt;/p&gt;

&lt;p&gt;The similarities lie deep in the structure. Both eras reached the exhausted final stage of a tournament for advancement. The Han system of classics and recommendation and today’s educational arms race plus civil-service examinations both demand enormous personal investment for a limited number of upward openings; both saw marginal returns plunge after operating for generations. Both eras suffered a severe credibility deficit in official discourse. The hypocrisy of &lt;em&gt;mingjiao&lt;/em&gt;, like the gulf today between “positive energy” rhetoric and lived experience, forces the elite to maintain an increasingly expensive firewall between public expression and private judgement. And in both eras, the elite drew energy back from public participation into personal aesthetics and interiority. The drinking, music, and “pure conversation” of the Seven Sages of the Bamboo Grove are functionally akin to temple trips, tarot, podcasts, and audiobooks today.&lt;/p&gt;

&lt;p&gt;Liu Yiqing’s &lt;em&gt;A New Account of the Tales of the World&lt;/em&gt; records a detail that could be pasted directly into contemporary social media. Wang Ziyou, or Wang Huizhi, was living in Shanyin. He awoke one snowy night, opened the window, drank wine, and suddenly thought of his friend Dai Andao, who lived by the Shan River. He took a boat through the night. After travelling until morning and reaching Dai’s door, he turned around without going in. Asked why, he replied:&lt;/p&gt;

&lt;blockquote&gt;
  &lt;p&gt;“I went on a tide of feeling and returned when the feeling was spent. Why must I see Dai?”&lt;sup id=&quot;fnref:4-en&quot; role=&quot;doc-noteref&quot;&gt;&lt;a href=&quot;#fn:4-en&quot; class=&quot;footnote&quot; rel=&quot;footnote&quot;&gt;4&lt;/a&gt;&lt;/sup&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;“Going on a tide of feeling and returning when it is spent” miniaturises the entire Wei-Jin aesthetic: the process is the purpose, the experience the meaning, and only an act that serves no result is free. Young people today call this “listening to your heart” or “being yourself.” Their whirlwind travel, City Walks, and activity-partner relationships practise a reduced form of the same ideal. The difference is that Wang could follow his feeling because the estates of the Wang clan of Langya stood behind him. For a young person today, the feeling often ends abruptly when the credit-card bill arrives.&lt;/p&gt;

&lt;p&gt;The differences matter just as much. Ignore them, and the analogy collapses into lyricism.&lt;/p&gt;

&lt;p&gt;The first difference is the availability of retreat. The great Wei-Jin clans owned self-sufficient estates; “retiring to the mountains” was a genuine material option. Tao Yuanming could “pick chrysanthemums beneath the eastern fence” because he had an eastern fence. Young people today have no estates. They have a ten-square-metre rented room and food delivered in thirty minutes. The division of labour in modern society permits no genuine exit. You may lie flat spiritually, but every month rent, social insurance, and phone bills demand that you connect to the system somehow. &lt;mark&gt;Exit is gestural; connection is structural.&lt;/mark&gt; Contemporary “lying flat” is therefore more psychologically torturous than Wei-Jin reclusion. It is an incomplete withdrawal, a permanent oscillation between participation and refusal.&lt;/p&gt;

&lt;p&gt;The second difference is the consequence. Three centuries of division and warfare followed Wei-Jin “pure conversation”—the Sixteen Kingdoms and the Northern and Southern dynasties, accompanied by enormous population loss and partial ruptures in civilisation. Many who invoke the Wei-Jin analogy point implicitly toward this terrifying ending: spiritual anomie heralds political collapse. Methodologically, however, the inference does not stand. The political breakdown had independent and sufficient structural causes: the rise of military aristocracies, a vacuum in central military power, and migration pressures from northern peoples. No rigorous causal relationship between pure conversation and collapse has ever been established. &lt;strong&gt;Spiritual anomie may precede collapse, or it may simply be a society’s adjustment under pressure.&lt;/strong&gt; Countless intermediate links requiring independent proof lie between it and political disintegration.&lt;/p&gt;

&lt;p&gt;Where, then, is today’s spiritual anomie more likely to lead?&lt;/p&gt;

&lt;h2 id=&quot;a-future-closer-to-japan&quot;&gt;A future closer to Japan&lt;/h2&gt;

&lt;p&gt;A more recent and comparable reference is Japan after the 1990s.&lt;/p&gt;

&lt;p&gt;In his 2015 book &lt;em&gt;The Low-Desire Society&lt;/em&gt;, Kenichi Ohmae offered an almost anthropological account of young Japanese people after the bubble burst:&lt;/p&gt;

&lt;blockquote&gt;
  &lt;p&gt;“Young Japanese have no desire, no dreams, no drive. … They do not want to rise above others, and do not want cars or luxury goods. Their willingness to consume has fallen to its lowest point.”&lt;sup id=&quot;fnref:5-en&quot; role=&quot;doc-noteref&quot;&gt;&lt;a href=&quot;#fn:5-en&quot; class=&quot;footnote&quot; rel=&quot;footnote&quot;&gt;5&lt;/a&gt;&lt;/sup&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Ohmae named this condition the “low-desire society.” Its features include young people refusing mortgages, delaying or forgoing marriage and children, reducing consumption to daily necessities, losing interest in promotion, and turning to small private satisfactions. Applied to urban Chinese youth in 2024, the list needs hardly a word changed.&lt;/p&gt;

&lt;p&gt;Japan offers a decisive empirical counterexample: spiritual anomie does not necessarily lead to political collapse. In the “lost thirty years” after the bubble burst, GDP growth hovered around zero and the population continued to shrink, but social order remained stable, democratic institutions continued to operate, and the crime rate even declined. A low-desire society can be quiet, orderly, and slowly ageing. Japan demonstrated one possible modernity: a generation’s disappointment need not end in violent rupture; it can be absorbed through prolonged stagnation and gentle contraction.&lt;/p&gt;

&lt;p&gt;Through this framework, China’s spiritual trajectory in the 2020s is more likely to resemble Japanese low desire than Wei-Jin dynastic disintegration. Involution, lying flat, offering incense, and tapping a digital wooden fish all retreat rather than confront; they lower expectations rather than overturn the table. They signal the arrival of modernity, not the end of civilisation. When a society is prosperous, safe, and atomised enough that individuals can reject mainstream values alone, without collective action, the rejection naturally tends toward quiet rather than violence.&lt;/p&gt;

&lt;p&gt;Yet the Japanese analogy also requires discipline. Three fissures separate China and Japan. First, Japan grew old with per-capita GDP above forty thousand US dollars, a mature social safety net, and universal health and pension coverage. China is growing old at roughly thirteen thousand dollars per capita, and nearly 300 million migrant workers lack urban employee pension insurance. Low desire lacks the material floor it had in Japan. Second, Japan stagnated under the American security umbrella with trade channels broadly open; China would stagnate behind tariff walls and technological restrictions. Third and most fundamentally, Japan has elections as a pressure-release valve. Prime ministers changed more than a dozen times in three decades, and policy shifted with public opinion. Disappointment could be expressed through voting, however limited the effect. China has no such buffer. When disappointment can be released neither through exit—the social structure prevents genuine withdrawal—nor through voice—the political structure provides no channel for it—it can only settle. Sedimented disappointment does not disappear. It merely becomes silent and dense, waiting for an unpredictable moment to find an unpredictable outlet.&lt;/p&gt;

&lt;p&gt;Albert Hirschman’s classic 1970 framework has an almost cruel explanatory force here. Faced with a disappointing organisation or system, the individual has three options: &lt;strong&gt;exit, voice, or loyalty&lt;/strong&gt;. Hirschman wrote:&lt;/p&gt;

&lt;blockquote&gt;
  &lt;p&gt;“The exit option is widely held to be uniquely important in keeping sellers alert. It is the threat of exit that acts as the primary restraining force.”&lt;sup id=&quot;fnref:6-en&quot; role=&quot;doc-noteref&quot;&gt;&lt;a href=&quot;#fn:6-en&quot; class=&quot;footnote&quot; rel=&quot;footnote&quot;&gt;6&lt;/a&gt;&lt;/sup&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The threat of exit is the chief force compelling an organisation to correct itself. But in a society where geographic exit is costly, political voice is narrow, and loyalty is exhausted, all three roads are half blocked. A fourth option appears, unnamed in Hirschman’s model: remain in place, but turn the volume all the way down. Do not leave, shout, or pledge loyalty; quietly lower every expectation. That is what is happening today. &lt;mark&gt;It is quieter than exit, safer than voice, and more honest than loyalty.&lt;/mark&gt;&lt;/p&gt;

&lt;h2 id=&quot;the-disappearance-of-the-nearby&quot;&gt;The disappearance of “the nearby”&lt;/h2&gt;

&lt;p&gt;Of all analyses of contemporary China’s spiritual condition, one concept proposed by the anthropologist Xiang Biao is particularly penetrating: &lt;strong&gt;the disappearance of the nearby&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;In a conversation with Xu Zhiyuan, Xiang described the peculiar dumbbell shape of contemporary Chinese life. At one end sits the intensely private self—the phone screen and the inner world. At the other sits the grand collective imaginary—the nation, the people, and the global order. The level between them, “the nearby,” is systematically disappearing: the neighbourhood where you live, your neighbours, the streets you walk each day, and the concrete people and places with which you interact face to face. He summarised it this way:&lt;/p&gt;

&lt;blockquote&gt;
  &lt;p&gt;“There is a major problem in how our attention is distributed. We pay great attention to the self and to the world, but the intermediate level of ‘the nearby,’ as a basic unit through which life is perceived, has been hollowed out.”&lt;sup id=&quot;fnref:7-en&quot; role=&quot;doc-noteref&quot;&gt;&lt;a href=&quot;#fn:7-en&quot; class=&quot;footnote&quot; rel=&quot;footnote&quot;&gt;7&lt;/a&gt;&lt;/sup&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;&lt;strong&gt;The disappearance of the nearby is not a spatial problem. It is a problem of meaning.&lt;/strong&gt; When you do not know your neighbours, care about changes on your street, or participate in collective life at the community level, the most basic intermediary between you and the world is missing. No matter how stirring grand narratives such as “national rejuvenation” and “the rise of a great power” may be, they cannot directly redeem the individual’s need for everyday meaning. No matter how refined a purely private life of consumption, entertainment, and self-care becomes, it lacks recognition and response from others. &lt;mark&gt;People need the nearby because it is the soil in which meaning grows: in the neighbourhood, you are known, needed, and remembered, and your actions produce visible consequences within reach.&lt;/mark&gt;&lt;/p&gt;

&lt;p&gt;Xiang’s diagnosis resonates deeply with Durkheim’s concept of anomie. Durkheim traced modern anomie to the breakdown of traditional community. Xiang identifies the precise level at which the breakdown occurs: not the most intimate family, which remains, nor the most distant state, which is everywhere, but the middle layer through which the individual forms an organic bond with society.&lt;/p&gt;

&lt;p&gt;In China, several forces accelerate the disappearance. Urbanisation uprooted hundreds of millions from communities of acquaintances and replanted them in compounds of strangers. Platform economies moved every daily need—food, travel, shopping, and social life—from the street to the screen, stripping streets of their function as sites of public life. Community-level self-government has long been institutionally absent, with residents’ committees operating more often as administrative terminals than vehicles of resident autonomy. The nearby has become like a sheet pulled from above and below, stretched thinner and thinner until it turns transparent.&lt;/p&gt;

&lt;p&gt;Suspension is thus the existential background of this generation. You can shed tears over the maiden flight of a Chinese-built airliner on your phone, then turn around unable to name the neighbour opposite. You can know every detail of the China-US chip contest but not know who is responsible for the pipes on your floor. Grand narratives offer pride and the private world offers comfort, but no bridge connects them. The psychological cost of suspension is a diffuse unreality: everything operates and you operate within it, yet you cannot determine your relationship to the machine. You are not one of its parts—you can be replaced at any time—but neither are you its owner—you have no power of decision. You merely happen to be present.&lt;/p&gt;

&lt;h2 id=&quot;a-silent-refusal&quot;&gt;A silent refusal&lt;/h2&gt;

&lt;p&gt;On the surface, the new pure conversation is a consumer phenomenon: temple tickets, tarot sales, MBTI labels, and downloads of the digital wooden fish. At the middle level, it is a social-psychological phenomenon: after returns from the orthodox route collapse, a generation withdraws its energy from external competition into inner experience, and from long-term planning into immediate sensation. At the deepest level lies a structural crisis of meaning: the old social contract—work hard and you will rise—has failed, while no new contract has formed. The nearby has been hollowed out. Individuals float in the vacuum between private life and the state, using fragments of spiritual consumption to maintain a minimum supply of meaning.&lt;/p&gt;

&lt;p&gt;The Wei-Jin analogy illuminates half the situation. When official discourse loses credibility and the marginal return from the tournament of advancement falls to zero, the elite’s spiritual turn can display a structural similarity across seventeen centuries. But the analogy must stop there. The three centuries of division after Wei-Jin had independent military and political causes and cannot be deduced directly from spiritual anomie. Japan is a more plausible reference: the slow ageing of a low-desire society—quiet and orderly, but also dull and stagnant. Yet the three fissures between China and Japan—income, external conditions, and pressure-release mechanisms—mean that a Chinese low-desire society would be rougher, tenser, and less cushioned.&lt;/p&gt;

&lt;p&gt;Before his execution, Ji Kang finished his final piece and said, “From this day, &lt;em&gt;Guangling San&lt;/em&gt; is lost.”&lt;sup id=&quot;fnref:8-en&quot; role=&quot;doc-noteref&quot;&gt;&lt;a href=&quot;#fn:8-en&quot; class=&quot;footnote&quot; rel=&quot;footnote&quot;&gt;8&lt;/a&gt;&lt;/sup&gt; To him, the loss of a composition was more regrettable than the end of a life. Such an extreme elevation of aesthetics above survival cannot be reproduced today. But his central act—the refusal expressed through silence, non-cooperation, and turning away—is being reproduced by the hundred million. The difference is that Ji Kang knew what he rejected. Many young people today cannot necessarily articulate what they reject. They are simply tired: tired of the track, tired of performance, tired of the promise that “you only need to work a little harder.”&lt;/p&gt;

&lt;p&gt;And so they offer incense. Merit +1.&lt;/p&gt;

&lt;hr /&gt;

&lt;p&gt;&lt;strong&gt;Notes&lt;/strong&gt;&lt;/p&gt;

&lt;div class=&quot;footnotes&quot; role=&quot;doc-endnotes&quot;&gt;
  &lt;ol&gt;
    &lt;li id=&quot;fn:1-en&quot; role=&quot;doc-endnote&quot;&gt;
      &lt;p&gt;Émile Durkheim, &lt;em&gt;Suicide: A Study in Sociology&lt;/em&gt;, trans. John A. Spaulding and George Simpson, The Free Press, 1951 (originally published 1897), p. 258. &lt;a href=&quot;#fnref:1-en&quot; class=&quot;reversefootnote&quot; role=&quot;doc-backlink&quot;&gt;&amp;#8617;&lt;/a&gt;&lt;/p&gt;
    &lt;/li&gt;
    &lt;li id=&quot;fn:2-en&quot; role=&quot;doc-endnote&quot;&gt;
      &lt;p&gt;Yu Yingshi, &lt;em&gt;The Scholar and Chinese Culture&lt;/em&gt;, Shanghai People’s Publishing House, 2003, p. 335. &lt;a href=&quot;#fnref:2-en&quot; class=&quot;reversefootnote&quot; role=&quot;doc-backlink&quot;&gt;&amp;#8617;&lt;/a&gt;&lt;/p&gt;
    &lt;/li&gt;
    &lt;li id=&quot;fn:3-en&quot; role=&quot;doc-endnote&quot;&gt;
      &lt;p&gt;Ji Kang, &lt;em&gt;Letter Breaking Off Relations with Shan Juyuan&lt;/em&gt;; see &lt;em&gt;Book of Jin: Biography of Ji Kang&lt;/em&gt;. &lt;a href=&quot;#fnref:3-en&quot; class=&quot;reversefootnote&quot; role=&quot;doc-backlink&quot;&gt;&amp;#8617;&lt;/a&gt;&lt;/p&gt;
    &lt;/li&gt;
    &lt;li id=&quot;fn:4-en&quot; role=&quot;doc-endnote&quot;&gt;
      &lt;p&gt;Liu Yiqing, &lt;em&gt;A New Account of the Tales of the World&lt;/em&gt;, “Uninhibitedness.” &lt;a href=&quot;#fnref:4-en&quot; class=&quot;reversefootnote&quot; role=&quot;doc-backlink&quot;&gt;&amp;#8617;&lt;/a&gt;&lt;/p&gt;
    &lt;/li&gt;
    &lt;li id=&quot;fn:5-en&quot; role=&quot;doc-endnote&quot;&gt;
      &lt;p&gt;Kenichi Ohmae, &lt;em&gt;The Low-Desire Society: A New Theory of National Wealth for an Age Without Great Ambitions&lt;/em&gt;, Shogakukan Shinsho, 2015. &lt;a href=&quot;#fnref:5-en&quot; class=&quot;reversefootnote&quot; role=&quot;doc-backlink&quot;&gt;&amp;#8617;&lt;/a&gt;&lt;/p&gt;
    &lt;/li&gt;
    &lt;li id=&quot;fn:6-en&quot; role=&quot;doc-endnote&quot;&gt;
      &lt;p&gt;Albert O. Hirschman, &lt;em&gt;Exit, Voice, and Loyalty: Responses to Decline in Firms, Organizations, and States&lt;/em&gt;, Harvard University Press, 1970, p. 21. &lt;a href=&quot;#fnref:6-en&quot; class=&quot;reversefootnote&quot; role=&quot;doc-backlink&quot;&gt;&amp;#8617;&lt;/a&gt;&lt;/p&gt;
    &lt;/li&gt;
    &lt;li id=&quot;fn:7-en&quot; role=&quot;doc-endnote&quot;&gt;
      &lt;p&gt;Xiang Biao, in conversation with Xu Zhiyuan, &lt;em&gt;Thirteen Invitations&lt;/em&gt;, season four, 2019. Paraphrased here. &lt;a href=&quot;#fnref:7-en&quot; class=&quot;reversefootnote&quot; role=&quot;doc-backlink&quot;&gt;&amp;#8617;&lt;/a&gt;&lt;/p&gt;
    &lt;/li&gt;
    &lt;li id=&quot;fn:8-en&quot; role=&quot;doc-endnote&quot;&gt;
      &lt;p&gt;Ji Kang’s words before execution; see &lt;em&gt;A New Account of the Tales of the World&lt;/em&gt;, “Cultivated Tolerance.” &lt;a href=&quot;#fnref:8-en&quot; class=&quot;reversefootnote&quot; role=&quot;doc-backlink&quot;&gt;&amp;#8617;&lt;/a&gt;&lt;/p&gt;
    &lt;/li&gt;
  &lt;/ol&gt;
&lt;/div&gt;
</content>
  </entry>
  
  <entry>
    <title>通用智能时代，我们将何去何从</title>
    <link href="https://mochiaochen.github.io/writing/2026/07/general-intelligence-age/" rel="alternate" type="text/html"/>
    <published>2026-07-22T16:30:00+08:00</published>
    <updated>2026-07-22T16:30:00+08:00</updated>
    <id>https://mochiaochen.github.io/writing/2026/07/general-intelligence-age</id>
    <content type="html" xml:base="https://mochiaochen.github.io/writing/2026/07/general-intelligence-age/">&lt;h2 id=&quot;引言我们正站在什么样的历史节点上&quot;&gt;引言：我们正站在什么样的历史节点上&lt;/h2&gt;
&lt;p&gt;2025年，人工智能的发展进入了一个前所未有的临界点。如果说2023年ChatGPT的爆发让全球意识到大语言模型（LLM）的知识涌现能力，那么2025年的核心叙事已经转向了一个更深刻的命题：能动性（Agency）的边际成本正在趋零。&lt;/p&gt;

&lt;p&gt;回顾信息技术的演进史，我们可以识别出三次边际成本归零的结构性拐点。1995年前后，互联网的普及使信息获取的边际成本急速坍缩，信息开始变得无处不在。2023年，以GPT系列为代表的大语言模型将公开知识的获取成本压至近零，任何人都可以通过自然语言对话获取专业级的知识解答。而2025年，随着智能体（Agent）生态系统的初步成型，我们正在见证第三次拐点的到来：能动性本身开始变得可规模化供给。所谓「能动性」，指的是系统自主理解意图、制定计划、调用工具、执行任务并根据反馈进行调整的能力。当这种能力的部署成本持续下降，人类社会的生产力结构将发生根本性的重塑。&lt;/p&gt;

&lt;p&gt;要理解这一变革的全貌，我们需要一个全栈式的分析框架。本文将沿着三个维度展开：战略格局、发展趋势、研发模式。这三个维度并非平行排列，它们构成一个从宏观到微观、从「为什么」到「怎么做」的递进体系。&lt;/p&gt;

&lt;h2 id=&quot;第一部分战略格局通用智能时代的文明博弈&quot;&gt;第一部分：战略格局——通用智能时代的文明博弈&lt;/h2&gt;
&lt;h3 id=&quot;一全栈结构理解ai产业的系统性框架&quot;&gt;一、全栈结构：理解AI产业的系统性框架&lt;/h3&gt;
&lt;p&gt;理解人工智能的战略格局，首先需要一张「全栈地图」。这张地图从底层到顶层依次包含：&lt;/p&gt;

&lt;p&gt;基础设施层涵盖能源功耗、矿业资源、材料科学、设备制造。AI的底层依赖是物理性的。训练一个前沿大模型消耗的电力以吉瓦时计，而推理阶段的持续运行同样需要庞大的能源供给。这就引出了一个关键的地缘约束：算力基础设施的选址和建设，本质上受制于能源供给和地理环境。美国推进StarGate计划，其核心逻辑正是将算力基建与能源基建捆绑，形成「算力-能源」综合体。中国则依托光伏产业的全球领先地位和核能技术的持续推进（包括小型模块化反应堆和可控核聚变的探索），试图在能源维度构建差异化优势。&lt;/p&gt;

&lt;p&gt;智算体系层包含智算中心、智能云、数据中心等核心设施。这一层的竞争焦点正在从单纯的算力规模转向「算力-模型协同设计」（Co-Design）的效率。传统的云计算范式以通用CPU为核心，而智算时代要求以GPU张量核心为基础的浮点矩阵运算能力，并进一步向多芯片异构架构演进。&lt;/p&gt;

&lt;p&gt;模型训练层是当前竞争最为激烈的阵地，涵盖认知智能、情景智能、具身智能、空间智能和科学智能五大分支。这五种智能形态并非孤立的技术方向，它们共同构成通用智能的能力光谱。&lt;/p&gt;

&lt;p&gt;智能外延层则指向智能体、软件3.0体系以及各类终端应用与设备。这一层是AI能力向现实世界「扩散」的界面。&lt;/p&gt;

&lt;p&gt;产业结构层和扩散体系层关注的是AI如何重塑市场结构、如何向不同人群和场景渗透。2C、2B、2G三个市场维度各有其扩散逻辑和时间节奏。&lt;/p&gt;

&lt;p&gt;这个全栈结构的意义在于：它提醒我们，AI的竞争从来不是单一技术维度的较量。一个国家或地区的AI实力取决于其在全栈每一层的能力积累与层间的耦合效率。&lt;/p&gt;

&lt;h3 id=&quot;二新源头谁在驱动前沿&quot;&gt;二、新源头：谁在驱动前沿？&lt;/h3&gt;
&lt;p&gt;在全栈结构的顶端，有一个关键的驱动要素需要被特别关注：头部驱动力量的组织形态。&lt;/p&gt;

&lt;p&gt;中美两国在这一维度上呈现出鲜明的结构性差异。美国的前沿智能研发主要由研究型闭源创业公司（Research-Intensive Startups）驱动。OpenAI、Anthropic、xAI、Thinking Machines Lab、SSI等公司构成了一个独特的创新群落，它们兼具学术研究的深度和工程落地的速度。这些公司的共同特征是：由顶级AI研究者创办或主导，将基础研究与产品工程紧密耦合，并通过巨额融资维持高强度的「研究密度」。头部大厂如NVIDIA和Google则以基础设施和平台生态的身份参与竞争。&lt;/p&gt;

&lt;p&gt;中国的格局有所不同。字节跳动（豆包）、腾讯、阿里巴巴等大厂主导了大模型的研发和部署，开源生态（以DeepSeek为代表）则构成了重要的补充力量。中国在开源模型方面的贡献已经引起全球关注，DeepSeek R1的发布标志着中国在推理能力方面取得了令人瞩目的进展。然而，纯粹的研究型创业公司在中国的生态中相对稀缺，这与人才体系、资本生态和组织文化等深层因素有关。&lt;/p&gt;

&lt;p&gt;一个值得深入讨论的概念是Researcher Founder，即研究者型创始人。在科学第四范式（计算驱动的科学）的背景下，从基础研究到产业应用的路径正在被极大缩短。传统的「产-学-研」链条假设了一个从大学基础研究到企业应用开发的线性传导过程。但在AI时代，研究型创业公司打破了这一线性假设：同一个组织同时承担基础研究、工程实现和价值创造的三重角色。OpenAI从一个非营利研究机构演变为全球估值最高的AI公司，正是这一模式最具戏剧性的例证。&lt;/p&gt;

&lt;p&gt;要支撑这种模式的可持续发展，至少需要四个条件：充足的人才体系、成熟的资本生态、规模化的算力资源，以及从研究到产品的闭环路径。OpenAI和Anthropic之所以能够持续引领前沿，正是因为它们同时满足了这四个条件。&lt;/p&gt;

&lt;h3 id=&quot;三新全球化ai扩散的地缘逻辑&quot;&gt;三、新全球化：AI扩散的地缘逻辑&lt;/h3&gt;
&lt;p&gt;AI正在重构全球化的底层逻辑。美国的AI扩散策略具有鲜明的全球推广特征：将美国的技术栈（从芯片架构到模型体系到应用生态）作为一个整体向全球输出，目前已占据全球AI技术栈约80%的份额。这一策略的战略意图是明确的——通过技术标准的锁定实现长期的生态控制。与此同时，美国还推出了Genesis Mission计划，试图用AI加速科学研究的进程，尤其是在核能、生物技术、先进材料等战略性领域。&lt;/p&gt;

&lt;p&gt;中国的对策则体现在「AI+」产业政策和「新数字丝绸之路」的构想中。「十五五」规划将人工智能的战略地位进一步提升，而「新一带一路」的数字化版本则试图在全球南方国家建立中国AI技术栈的应用基础。这种竞争本质上是两种技术文明扩散路径的博弈。&lt;/p&gt;

&lt;p&gt;值得注意的是，这种博弈的底层变量正在发生变化。传统的技术霸权依赖于闭源专有系统的控制力，而开源运动正在动摇这一基础。当DeepSeek以开源方式发布高性能推理模型时，它同时在技术竞争和地缘政治两个维度上产生了冲击：它证明了闭源壁垒可以被绕过，也为全球开发者提供了一个不依赖美国技术栈的替代路径。&lt;/p&gt;

&lt;h3 id=&quot;四新基建80年周期律与泡沫风险&quot;&gt;四、新基建：80年周期律与泡沫风险&lt;/h3&gt;
&lt;p&gt;从历史规律的视角看，重大技术革命的基础设施建设阶段往往伴随着大规模的资本投入、周期性的泡沫，以及泡沫破裂后的「新黎明」。铁路、电力、互联网，无一例外。&lt;/p&gt;

&lt;p&gt;当前的AI基建浪潮同样遵循着类似的节奏。美国的StarGate计划代表了一种「基建经济」的思路：通过万亿级的算力和能源投资，抢占未来十年AI基础设施的制高点。但这种高度集中的资本投入也蕴含着泡沫风险，尤其是当AI应用的商业化收入在短期内无法匹配基建投入的规模时。&lt;/p&gt;

&lt;p&gt;Carlota Perez在其经典著作《技术革命与金融资本》中提出了技术革命的「四阶段模型」：爆发期、狂热期、协同期和成熟期。狂热期（Frenzy）的典型特征是金融资本过度涌入新技术领域，推高资产价格，形成泡沫。泡沫的破裂并不意味着技术本身失败，它更类似一次「强制性的资源重新配置」，为后续的协同期（Synergy）奠定基础。如果我们将当前的AI投资热潮置于这一框架中，有理由认为我们正处于从爆发期向狂热期过渡的阶段。&lt;/p&gt;

&lt;p&gt;中国的策略选择与美国形成了对照。政府引导、平衡发展的模式意味着：由国家和地方政府主导智算中心（训练场、实训场）的建设，避免纯粹市场力量驱动下的过度投资。在能源维度，中国的光伏产业已具备全球压倒性优势，核能技术（小型反应堆、可控聚变探索）也在持续推进，这为AI基建的能源需求提供了长期保障。&lt;/p&gt;

&lt;p&gt;人类社会大约以80年为周期经历一次范式级的技术革命。如果这一周期律仍然有效，我们正站在一个新的80年周期的开端。&lt;/p&gt;

&lt;h3 id=&quot;五新数字化平台生态与拐点效应&quot;&gt;五、新数字化：平台生态与拐点效应&lt;/h3&gt;
&lt;p&gt;数字经济的核心规律是平台-生态驱动。每一个计算范式的更迭都伴随着新平台的崛起和围绕平台形成的生态系统的繁荣。PC时代的微软和英特尔、互联网时代的谷歌和亚马逊、移动互联网时代的苹果和微信，都验证了这一规律。&lt;/p&gt;

&lt;p&gt;在AI时代，新的平台形态正在成型。美国凭借其在前沿模型和云基础设施方面的先发优势，正在构建全球性的AI平台生态。中国则在C端（消费者端）拥有独特的优势：庞大的移动互联网用户基础为AI应用的快速扩散提供了天然的土壤。在B端（企业端），中国面临着一个有趣的「跃迁机会」：由于中国企业的数字化进程整体上落后于美国，反而有可能跳过传统SaaS的中间阶段，直接进入AI原生的企业服务范式。&lt;/p&gt;

&lt;h3 id=&quot;六新价值与新产业&quot;&gt;六、新价值与新产业&lt;/h3&gt;
&lt;p&gt;AI带来的产业重构远不止于效率提升。从宏观视角看，它正在催生全新的产业门类。中国的「十五五」规划已经将半导体、人工智能、智能制造、量子计算、低空经济等列为战略性新兴产业。这些产业之间存在深度的技术耦合：AI芯片的突破依赖半导体工艺的进步，智能制造需要具身智能的支撑，低空经济则是空间智能的典型应用场景。&lt;/p&gt;

&lt;p&gt;一个常被忽视但同样重要的维度是美国的再工业化需求。美国经济长期面临制造业空心化的结构性挑战，而AI（尤其是具身智能和工业机器人）被视为实现再工业化的关键技术杠杆。这一需求正在深刻影响美国AI政策的优先级排列。&lt;/p&gt;

&lt;h3 id=&quot;七文明体系科学范式的更迭&quot;&gt;七、文明体系：科学范式的更迭&lt;/h3&gt;
&lt;p&gt;将视野进一步拉升到文明尺度，AI的意义在于它正在推动科学研究范式的第四次更迭。&lt;/p&gt;

&lt;p&gt;科学史可以被粗略地划分为四个范式。第一范式是经验驱动的时代，以观察和思考为核心工具，知识以口传心授的方式极缓慢地扩散——从古埃及的灌溉规律到中国的二十四节气，人类用了数千年积累最初的经验体系。第二范式以伽利略的系统性实验为标志，引入了控制变量和可重复验证的方法论，知识扩散的速度缩短到了数十年的量级。第三范式是理论驱动的时代，牛顿力学、麦克斯韦方程组、爱因斯坦的相对论都是这一范式的巅峰之作，数学建模和同行评审构成了知识生产和验证的标准流程。&lt;/p&gt;

&lt;p&gt;第四范式由Jim Gray在2007年提出，其核心特征是计算驱动：算法模型、大规模数据采集、人机协同和迭代式验证。AlphaFold对蛋白质结构的预测、自动驾驶系统的端到端学习、脑机接口的信号解码，都是第四范式的典型产物。在第四范式中，科学发现的速度被极大加快——从年和月缩短到周和天，领域边界也被显著拓宽。&lt;/p&gt;

&lt;p&gt;这一范式更迭的深远意义在于它改变了科学创新的组织形态。在第三范式中，大学是创新的核心，科研成果通过技术转移机制（如美国的Bayh-Dole法案）进入产业。在第四范式中，研究型创业公司成为新的核心组织：它们同时拥有研究能力、工程能力和市场能力，能够在「从-1到1」的全过程中进行系统性的加速创新。OpenAI、DeepMind、DeepSeek都是这一新范式的典型代表。&lt;/p&gt;

&lt;p&gt;美国在1944年提出的《科学：无尽的前沿》（Science: The Endless Frontier）奠定了二战后美国科技创新体系的战略蓝图。近80年后的今天，AI正在驱动一次新的科技战略重构。美国的Genesis Mission瞄准先进制造、生物技术、能源技术（核能现代化、聚变技术）、电网技术等方向。中国则首次将人工智能提升到国家战略的最高层级，并着手探索新一代科研体系。&lt;/p&gt;

&lt;p&gt;这场「新产-学-研」组合的重构，本质上是第三范式向第四范式过渡过程中，创新要素的重新排列组合。&lt;/p&gt;

&lt;h3 id=&quot;八技术生产关系环境与人口&quot;&gt;八、技术、生产关系、环境与人口&lt;/h3&gt;
&lt;p&gt;全栈结构的文明体系维度还涵盖了几个同样关键的领域。&lt;/p&gt;

&lt;p&gt;在技术方面，AI正在加速四大前沿科技方向的突破：新能源科技、新生命科技、新材料科技和新空间科技。这四个方向的共同特征是：它们的研究过程高度依赖计算模拟、数据驱动和AI辅助的实验设计。例如，AI在材料科学中的应用已经能够从数百万种可能的化合物中快速筛选出具有特定性能的候选材料，将传统需要数年的发现过程压缩到数周。&lt;/p&gt;

&lt;p&gt;在生产关系方面，金融体系的演变正在成为AI时代的关键变量。发达的金融体系是AI创新生态的必要条件——从风险投资到IPO退出的完整链条支撑着研究型创业公司的成长。与此同时，美元体系的储备货币地位面临长期结构性压力，人民币的国际化进程和数字货币（包括稳定币）的发展正在改变全球资本流动的格局。AI本身也在深刻改变金融行业的运作方式，从量化交易到风险评估再到合规监控，智能化渗透已经无处不在。&lt;/p&gt;

&lt;p&gt;在环境维度，一个令人兴奋的前沿领域是空间拓展。从地面到低空、高空再到太空，人类活动空间的垂直延伸正在催生全新的产业生态。低空经济（无人机物流、城市空中交通、空域管理）正成为中国产业政策的重点方向之一。空天交通网、空域一体化、空天信息融合等概念正在从科幻走向工程实现。这里的关键技术支撑正是空间智能——AI对三维空间的理解、推理和规划能力。&lt;/p&gt;

&lt;p&gt;在人口方面，中国面临着深刻的结构性挑战。60岁以上人口已突破3.1亿，出生率降至6.77‰。这一人口结构变化同时催生了巨大的需求和约束：教育体系需要AI驱动的全面变革以适应新的人才培养需求；医疗健康领域则形成了万亿级的刚性需求，影像诊断、远程手术、智慧养老正在成为千亿级蓝海市场。从某种意义上说，AI是中国应对人口老龄化挑战的战略性工具。&lt;/p&gt;

&lt;h2 id=&quot;第二部分发展趋势从通用智能到智能体再到新数字化&quot;&gt;第二部分：发展趋势——从通用智能到智能体再到新数字化&lt;/h2&gt;
&lt;h3 id=&quot;一通用智能的发展阶段从学知识到学组织&quot;&gt;一、通用智能的发展阶段：从学知识到学组织&lt;/h3&gt;
&lt;p&gt;通用智能的发展可以被划分为五个层次递进的阶段，每个阶段对应着模型能力的质变跃迁。&lt;/p&gt;

&lt;p&gt;第一层：学知识。 这一阶段的核心技术是预训练（Pre-training），模型通过海量的互联网文本数据获取公开知识。以GPT-3为代表的早期大语言模型展示了令人惊讶的知识涌现能力，中国的豆包也在这一阶段取得了优异表现。ChatGPT的现象级爆发正是第一层能力成熟的标志：人们第一次可以通过自然语言对话获取几乎任何领域的知识。这一阶段的Scaling Law主要基于上下文长度和参数规模。&lt;/p&gt;

&lt;p&gt;第二层：学推理。 2024年以来的核心突破在于后训练（Post-training）技术对模型推理能力的大幅提升。OpenAI的o系列模型和中国DeepSeek R1展示了通过强化学习和自我反思（self-reflection）大幅增强模型逻辑推理和问题求解能力的可能性。DeepSeek的开源策略进一步加速了推理能力的全球扩散。Deep Research等半现象级产品的出现标志着第二层能力已经开始产生实用价值。这一阶段的Scaling Law基于自洽反思的迭代深度。&lt;/p&gt;

&lt;p&gt;第三层：学能动。 这是2025年的核心战场。能动性的培养需要续训练（Continual Training），使模型获得意图理解、指令遵循、长程规划和工具交互的综合能力。这一层又可以细分为两个方向：数字世界的能动性（以Claude Code为典型代表，覆盖产品设计、科研开发、营销运营等全链条工作流）和物理世界的能动性（机器人、自动驾驶等具身智能场景）。Anthropic估值一年内上涨20倍，正是市场对能动性价值最直接的定价。在数字世界方向，Claude Code的现象级成功证明了当模型具备了在真实开发环境中自主完成复杂任务的能力时，其创造的价值是非线性增长的。在物理世界方向，世界模型（World Model）的多种技术路径已经起步，为具身智能提供物理直觉的基础。这一阶段的Scaling Law基于环境-工具-教育的三位一体。&lt;/p&gt;

&lt;p&gt;第四层：学创新。 这一阶段目前仍处于早期探索中，其目标是使AI系统具备真正的发明创新能力，能够突破现有知识的边界。AI辅助AI开发、AI科学家（AI Scientist）等概念初步进入实验阶段。IMO金牌级数学推理模型的出现暗示了这一方向的巨大潜力。这一阶段的Scaling Law将基于创新时间的积累。&lt;/p&gt;

&lt;p&gt;第五层：学组织。 这是通用智能的终极形态：多智能体系统具备复杂决策、组织协同、战略规划的能力。目前学术界对多智能体协调（Multi-Agent Coordination）的研究还处于非常早期的阶段，但其长远意义不亚于单体智能的突破。当多个高度自主的智能体能够像人类组织一样分工协作时，我们将面临一个全新的社会组织形态。这一阶段的Scaling Law将基于环境适应的复杂度。&lt;/p&gt;

&lt;h3 id=&quot;二scaling-law的新路径从模型到智能体的进化&quot;&gt;二、Scaling Law的新路径：从模型到智能体的进化&lt;/h3&gt;
&lt;p&gt;Scaling Law在过去三年中经历了深刻的范式扩展。最初的Scaling Law（Kaplan et al., 2020; Hoffmann et al., 2022）主要关注预训练阶段的参数量、数据量和计算量的幂律关系。但随着模型能力的提升和应用场景的深化，新的Scaling维度正在被发现和开发。&lt;/p&gt;

&lt;p&gt;一个有启发性的类比是将AI的演进与人类文明的社会化进程对照。从生物进化中的基因遗传，到语言的出现、教育体系的建立、工业化、数字化直至智能化，每一步都代表着能力传承和放大效率的跃升。AI的Scaling路径同样可以沿着这条线索理解：从数据驱动的「遗传」（预训练），到推理驱动的「认知」（后训练），再到环境交互驱动的「能动」（续训练），最终到创新和组织层面的更高阶Scaling。&lt;/p&gt;

&lt;p&gt;一个关键洞察是：新Scaling进程的速度和路径，在很大程度上由数字化环境和基础设施的成熟度决定。这意味着智能体时代的Scaling竞争不仅取决于模型本身的能力，更取决于为模型提供的「环境」——工具、数据、驾驶舱、上下文工程——的质量和完备程度。&lt;/p&gt;

&lt;h3 id=&quot;三通用智能的技术前沿&quot;&gt;三、通用智能的技术前沿&lt;/h3&gt;
&lt;p&gt;在全新的发展体系中，通用智能的各个分支正在快速推进。&lt;/p&gt;

&lt;p&gt;模型基础方面，当前的前沿研究集中在几个方向：稀疏注意力机制（Sparse Attention）正在大幅提升模型的计算效率，使超长上下文的处理成为可能；理解与生成的统一架构（Unified Architecture）试图在一个模型中同时实现高质量的语言理解和多模态内容生成；扩散模型（Diffusion Model）和流匹配（Flow Matching）方法在图像、视频、音频生成方面展示了惊人的能力；长上下文和记忆机制的突破使模型能够处理百万token级别的输入窗口；评测体系也在从标准化基准测试向真实场景、能动性导向的评估方法转变；持续学习（Continual Learning）的研究则聚焦于如何让模型在不遗忘已有知识的前提下持续获取新能力。&lt;/p&gt;

&lt;p&gt;认知智能的核心追求是高难度问题的求解能力和深层反思能力。当前的前沿方向包括：AI辅助AI开发（使用AI系统来改进AI系统的架构和训练过程）、AI科学家（自主提出假设、设计实验、分析结果的端到端研究系统），以及认知与能动性的深度结合——使模型不仅能「想」而且能「做」。超智能（Superintelligence）虽然仍是一个遥远的目标，但多家前沿研究机构已经在进行相关的安全和对齐研究。&lt;/p&gt;

&lt;p&gt;情景智能（Contextual Intelligence）关注的是模型对真实世界场景的感知和理解能力。全模态融合（音频、视频、触觉等多种输入通道的统一处理）是核心技术挑战。深情境（Deep Context）的概念强调模型需要理解的不仅是语言文本，还包括任务（Task）和奖励（Reward）的完整语义结构。随着智能体需要在越来越复杂的真实环境中运行，情景智能对原生数据采集的需求正在爆发式增长。&lt;/p&gt;

&lt;p&gt;具身智能（Embodied Intelligence）的进展近一年来显著加速。人形机器人（如Figure、Unitree等）、四足机器人（面向教育、娱乐和工业巡检）、工业机器人（操作、规划、行动）和自动驾驶构成了具身智能的四大应用场景。核心技术路径正在从传统的分模块架构向端到端（End-to-End）架构演进，世界模型（World Model）和VLA（Vision-Language-Action）模型成为关键突破方向。Physical Intelligence的Pi-0.5等探索性成果表明，将语言模型的推理能力迁移到物理操作领域是可行的。数据采集（本体数据+流程数据+标注数据）仍然是具身智能大规模落地的关键瓶颈。&lt;/p&gt;

&lt;p&gt;空间智能是一个正在快速浮现的方向。世界模型在空间维度的应用可以分为两类：数字世界的空间智能（以World Labs / Marble为代表，构建数字化的三维世界模型）和物理世界的空间智能（以奇岱松等为代表，专注于物理空间的上下文理解和空间智能机的开发）。空间推理能力和注意力机制的突破将直接加速具身智能在真实物理环境中的落地。&lt;/p&gt;

&lt;p&gt;科学智能（AI for Science, AI4S）正在成为AI最具战略价值的应用方向之一。其核心路径包括：建立面向科学研究的专属数据体系和模型体系；开发能够端到端参与科研过程的智能体系统；培养能够自主提出假设和验证假设的AI科学家。Genesis Mission在美国层面代表了这一方向的国家级战略投入。科学智能的终极愿景是用AI加速整个科学发现的过程——从数据采集、假设生成到实验验证，每一步都可以被AI显著加速。&lt;/p&gt;

&lt;h3 id=&quot;四智能体能动性的全面释放&quot;&gt;四、智能体：能动性的全面释放&lt;/h3&gt;
&lt;p&gt;2025年被广泛视为「智能体元年」。智能体的核心能力可以被分解为四个维度：意图理解（准确把握用户的真实需求）、指令遵循（严格按照指定要求执行任务）、长程规划（将复杂目标分解为可执行的步骤序列），以及工具交互（调用外部API、操作软件界面、读写文件系统）。&lt;/p&gt;

&lt;p&gt;但能动性的真正释放需要三个关键条件的支撑：&lt;/p&gt;

&lt;p&gt;真实环境。 智能体必须在真实场景中通过实践学习。从互联网上的静态文本到高价值真实场景的完整上下文，数据来源的转变是能动性突破的关键。智能体需要独立完成多轮、长期的复杂任务，在与真实环境的交互中持续提升规划和操作能力。&lt;/p&gt;

&lt;p&gt;真实评估。 传统的标准化基准测试（Benchmark）已经无法有效衡量智能体的真实能力。新一代评测体系强调：输入和输出必须来源于线上真实数据，确保评估的是真实用户需求；评测需要动态更新，定期引入最新的用户查询（Query），以避免模型对评测集的过拟合（Reward Hacking）。&lt;/p&gt;

&lt;p&gt;真实数据。 以「环境上下文」为核心的数据采集正在成为新的趋势。从泛化的互联网上下文转向高价值真实场景的完整、全面、多维度上下文，是智能体从「能用」到「好用」的关键转变。&lt;/p&gt;

&lt;h3 id=&quot;五驾驶舱人类掌控能动性的界面&quot;&gt;五、驾驶舱：人类掌控能动性的界面&lt;/h3&gt;
&lt;p&gt;「驾驶舱」（Cockpit）是一个在陆奇的框架中被反复强调的概念，它指的是人类利用和控制智能体能动性的界面系统。&lt;/p&gt;

&lt;p&gt;一个深刻的历史类比有助于理解这一概念的本质意义。内燃机的发明赋予了人类强大的机械动力（一种「物理世界的能动性」），但这种能力的释放经历了漫长的驾驶舱迭代过程：从马到马车到福特Model T到高速公路系统，每一次驾驶舱的升级都大幅扩展了机械动力的可用范围和使用效率。类似地，互联网赋予了人类信息处理的能动性，其驾驶舱从服务器发展到浏览器、搜索引擎、社交平台，每一代驾驶舱都释放了互联网能力的新维度。&lt;/p&gt;

&lt;p&gt;AI时代的能动性释放同样遵循这一规律。模型本身具备的能动性潜力是巨大的，但这种潜力需要经过驾驶舱的迭代和工程体系的构建才能被有效释放。OpenAI Codex是第一代AI编程驾驶舱，Claude Code代表了驾驶舱工程的重大进步，而围绕Claude Code正在形成的生态系统则预示着驾驶舱体系的进一步成熟。&lt;/p&gt;

&lt;p&gt;Claude Code的案例特别值得深入分析。其核心创新在于「硅基-碳基共享上下文」的理念。传统的云端AI方案（如早期的Codex和GitHub Copilot）要求开发者将工作环境迁移到云端，但现实中，大部分开发工作仍然在本地进行——本地文件系统、命令行、公司数据库、内部资源。Claude Code选择了一条截然不同的路径：作为一个本地化的命令行工具，它直接安装在开发者的PC上，完整访问本地环境。开发者无需改变工作习惯，而是通过目录级别的Markdown文件（用Prompt描述每个目录的环境和上下文），让AI能够「看见」并「使用」碳基世界的完整上下文。&lt;/p&gt;

&lt;p&gt;这种设计哲学的深层含义是：AI需要的上下文与人类的工作上下文应该是同一套。传统方案试图为AI构建一个独立的、纯数字化的运行环境，Claude Code的做法则是让AI融入人类已有的工作环境。这不仅降低了集成复杂度，更重要的是，它使得人与智能体之间的协作可以在共享的上下文中自然发生。&lt;/p&gt;

&lt;p&gt;驾驶舱的演进阶段可以类比自动驾驶的分级。在辅助驾驶阶段，人类主导、系统辅助，能动性较低；早期的对话模型（GPT-4、Claude 3.5）配合Coze、LangChain等工作流框架属于这一阶段。在准自动驾驶阶段，系统主导、人类接管，能动性处于中等水平；新一代能动推理模型（GPT-5、Claude 4.5级别）配合长程推理和环境交互能力，正在进入这一阶段。在全自动驾驶阶段，系统主导、零接管、高能动性，这是长远的发展方向。&lt;/p&gt;

&lt;p&gt;每个阶段都有其对应的工程挑战。辅助驾驶阶段需要工作流编排来弥补模型在多轮交互和长程推理方面的缺陷；准自动驾驶阶段的核心挑战是安全运行环境（沙盒、权限审计、行为护栏、人在回路）和上下文工程（数据语境化、工具语境化、上下文组织）；全自动驾驶阶段则需要高度成熟的交付质量保证体系（交付契约设计、评估Rubrics、自我检验、过程透明化）。&lt;/p&gt;

&lt;h3 id=&quot;六上下文工程智能体的操作系统&quot;&gt;六、上下文工程：智能体的操作系统&lt;/h3&gt;
&lt;p&gt;如果说驾驶舱是人类与智能体交互的「硬件」界面，那么上下文工程（Context Engineering）就是智能体运行的「操作系统」。&lt;/p&gt;

&lt;p&gt;上下文工程的核心体系可以分为三个层次：&lt;/p&gt;

&lt;p&gt;系统提示词（System Prompt） 承载组织性、目的性和个性化的上下文信息。它定义了智能体「是谁」、「为谁工作」、「遵循什么原则」。&lt;/p&gt;

&lt;p&gt;结构提示词（Structure Prompt） 承载静态的、结构性的环境上下文。类比物理办公空间，它定义了智能体的「工作空间」——哪些区域用于什么功能（会议区用于团队协作、办公区用于日常工作、洽谈区用于商务交流），以及组织结构中不同角色的职责分工。&lt;/p&gt;

&lt;p&gt;任务提示词（Task Prompt） 承载动态的任务上下文，从简单的工作流指令到复杂的「认知脚手架」（Cognitive Scaffolding）。认知脚手架是一个特别重要的概念：它不只是告诉智能体「做什么」，还提供了「怎么想」的框架——问题定义、解空间构建、认知维度分解、系统性思考过程等。&lt;/p&gt;

&lt;p&gt;上下文工程的完整体系还包括底层的支撑设施：驾驶舱体系（软件操作系统、文件系统+硬件设备）和上下文处理流程（感知环境→分析规划→任务执行→状态记忆→闭环提高）。&lt;/p&gt;

&lt;h3 id=&quot;七新数字化软件30的全面到来&quot;&gt;七、新数字化：软件3.0的全面到来&lt;/h3&gt;
&lt;p&gt;如果将软件的发展划分为三个范式，软件1.0是符号时代（代码输入、编译器执行）；软件2.0是数据时代（数据驱动、神经网络推理）；软件3.0则是语言时代（自然语言描述、大模型理解和执行）。&lt;/p&gt;

&lt;p&gt;软件3.0的技术栈从底层到顶层包含五个层次：&lt;/p&gt;

&lt;p&gt;硬件计算层：以GPU张量核心为基础的浮点矩阵运算，本质上是一种「分时计算时代」——计算资源通过时间片分配在不同的推理任务之间。&lt;/p&gt;

&lt;p&gt;系统核心层：通用操作系统由语言模型和世界模型构成，形成了一种全新的「上下文计算机」。在这个计算机中，Prompt是输入，模型推理是计算，Token生成是输出。&lt;/p&gt;

&lt;p&gt;商业化OS层：驾驶舱+上下文工程的组合构成了面向商业化的操作系统层。在物理世界的维度，空间驾驶舱+物理上下文则构成了「空间智能机」。&lt;/p&gt;

&lt;p&gt;拓展开发层：Plugins、Skills、Hooks、MCP（Model Context Protocol）等机制构成了类似于传统应用开发生态的拓展层。MCP尤其值得关注，它正在成为智能体接入外部工具和服务的标准协议。&lt;/p&gt;

&lt;p&gt;应用生态层：数字智能体（能动性、定制化、一次性APP）和具身智能体（机器人、自动驾驶、新设备终端）构成了面向终端用户的应用层。&lt;/p&gt;

&lt;h3 id=&quot;八产业结构的重塑&quot;&gt;八、产业结构的重塑&lt;/h3&gt;
&lt;p&gt;软件3.0对产业结构的影响是全方位的。&lt;/p&gt;

&lt;p&gt;在To C方向，AI正在重构全方位的流量入口。从通讯社交到内容游戏，从教育医疗到金融住房，几乎所有的消费场景都面临AI原生化的改造。这种改造不仅体现在用户界面的智能化，更体现在底层服务逻辑的重构——从标准化产品到个性化智能服务。&lt;/p&gt;

&lt;p&gt;在To B方向，新的产业结构机会正在形成。新IaaS（带动国产算力和训练体系的发展）、新PaaS（模型与算力协同，针对上下文提升效率）、新SaaS（面向能动性应用的新型服务体系）构成了企业级AI服务的三层架构。新的终端设备（车、机器人、无人机、无人船）和新的个人终端（AI眼镜、穿戴设备、手机）则构成了AI能力的物理延伸。&lt;/p&gt;

&lt;p&gt;从宏观视角看，一个类似于「AI工厂」的概念正在形成：企业的供应链、市场销售、客户支持、员工体验、金融法律、行政办公、政府关系等所有职能模块，都将通过AI驾驶舱连接到统一的智能操作系统上。&lt;/p&gt;

&lt;h2 id=&quot;第三部分研发模式从实验室到战场的落地路径&quot;&gt;第三部分：研发模式——从实验室到战场的落地路径&lt;/h2&gt;
&lt;h3 id=&quot;一新研发模式六个要素局势&quot;&gt;一、新研发模式：六个要素局势&lt;/h3&gt;
&lt;p&gt;新研发模式的核心骨架包含六个要素：算力体系（智能体云）、模型体系（能动性闭环）、数据体系（上下文工程）、驾驶舱模式（能动性释放）、新开发交付模式（高效定制共创）、新应用落地模式（系统性布局）。&lt;/p&gt;

&lt;h3 id=&quot;二新模型体系开源能动模型与评测生态&quot;&gt;二、新模型体系：开源能动模型与评测生态&lt;/h3&gt;
&lt;p&gt;在新模型体系的建设路径中，几个关键的建设方向值得关注：&lt;/p&gt;

&lt;p&gt;开源能动模型的开发。 在Qwen3、DeepSeek等开源基座模型的基础上，构建面向能动性的上层能力，包括意图理解、目标遵循、规划拆解、交互推理等核心能力模块。这些能力的泛化范围需要覆盖通用工具操作、网页交互、深度研究、代码编写等主要场景。&lt;/p&gt;

&lt;p&gt;真实场景评测系统（Open Agent Arena）。 传统的静态Benchmark已经无法满足智能体评估的需求。新的评测系统需要基于真实用户需求的动态评测，定期更新评测集以避免模型对评测的过拟合。&lt;/p&gt;

&lt;p&gt;训练即服务（Training as a Service）。 充分利用丰富的产业生态，将高价值真实场景与模型训练形成闭环：真实场景提供训练数据和评测标准，训练后的模型再部署到场景中验证效果，持续迭代。&lt;/p&gt;

&lt;p&gt;下一代能动性模型体系。 以NexAU等Agent开发框架为代表，构建端到端的Agentic数据合成管线和强化学习训练基础设施（NexRL、NexVenusCL），推动能动性模型从实验阶段走向产业级部署。&lt;/p&gt;

&lt;h3 id=&quot;三新算力体系从算力云到agent云&quot;&gt;三、新算力体系：从算力云到Agent云&lt;/h3&gt;
&lt;p&gt;算力体系的价值逻辑正在经历根本性的转变。&lt;/p&gt;

&lt;p&gt;旧模式以租卖算力（GPU时间片）为核心，利润空间有限，技术壁垒不高，面临严重的同质化竞争。升级到「Token as Service」（以模型推理为单位收费）虽然提升了价值密度，但仍然依赖公开互联网上下文，容易被替代。&lt;/p&gt;

&lt;p&gt;新模式的核心是Agent云——以智能体运行为中心的新型算力服务。这种模式的价值主张是「上下文即服务+智能体运行」：从服务人转向服务Agent，从互联网上下文转向高价值企业上下文，从按量计费转向以结果为导向的价值定价。这种转变的深层逻辑在于：企业的生产力上下文具有持续积累的特性——使用越久，沉淀的上下文越丰富，服务的价值越高。&lt;/p&gt;

&lt;p&gt;新算力体系的技术架构包含三层。新IaaS层提供智算云基础设施（先进芯片训练算力、国产替换、沙盒环境）。新PaaS层提供Agent能动基础设施（智能体运行时、上下文工程即服务、权限安全管理）。新SaaS层提供新企业生产力服务（智能体工程即服务、工作空间、能动配置、智能体驾驶舱）。&lt;/p&gt;

&lt;p&gt;贯穿这三层的是模型算力Co-Design的理念：多芯片异构算力体系与模型架构的协同优化，以及持续学习框架下的「训练即服务」（Training-as-a-Service）机制——自动调优、强化学习训练、异步并行机制等。&lt;/p&gt;

&lt;h3 id=&quot;四新数据体系从互联网数据到环境交互数据&quot;&gt;四、新数据体系：从互联网数据到环境交互数据&lt;/h3&gt;
&lt;p&gt;能动性时代的数据需求正在发生根本性的改变。&lt;/p&gt;

&lt;p&gt;预训练阶段的数据需求以互联网文本和书籍为主；后训练阶段转向对话数据和推理链数据；而能动性训练阶段则需要全新类型的数据——上下文数据和环境交互反馈数据。这些数据来源于真实的工作场景、业务流程和决策过程，具有高度的场景特异性和时效性。&lt;/p&gt;

&lt;p&gt;美国的模型公司和数据公司已经在系统性地构建面向职业体系的数据资产。几个案例值得关注：&lt;/p&gt;

&lt;p&gt;Anthropic和斯坦福大学都在利用美国劳工部的职业信息体系数据库（O*NET）研究AI对人类各种职业的影响，以此指引AI能动性的发展方向——本质上是在回答一个关键问题：「AI应该首先学会做哪些工作？」&lt;/p&gt;

&lt;p&gt;Surge AI正在进行大规模的企业环境构建，完整复现企业的工作软件环境、组织架构，虚构数万条企业数据并确保数据之间的逻辑一致性。这种「合成企业环境」为智能体的训练提供了安全、可控、可规模化的实践场所。&lt;/p&gt;

&lt;p&gt;OpenAI最新的「Mercury」项目则以时薪150美金的高价招募华尔街专业人士标注金融领域的数据，体现了高价值领域数据的稀缺性和获取难度。&lt;/p&gt;

&lt;p&gt;这些案例共同指向一个结论：能动性时代的数据竞争，核心在于「谁能构建出最完整、最真实、最高价值的场景上下文」。&lt;/p&gt;

&lt;h3 id=&quot;五驾驶舱体系的工程化&quot;&gt;五、驾驶舱体系的工程化&lt;/h3&gt;
&lt;p&gt;驾驶舱设计的核心原则是：利用好系统的能动性并弥补其缺陷。当前能动模型的主要缺陷包括：行为失控（在长程推理中偏离目标）、缺乏业务理解（对特定领域的知识和流程不熟悉）、交付质量波动（输出结果的稳定性不够）。&lt;/p&gt;

&lt;p&gt;针对这些缺陷，新驾驶舱体系构建了四层防护和增强机制：&lt;/p&gt;

&lt;p&gt;安全运行环境：安全运行沙盒、权限审计体系、行为护栏规则、人在回路（Human-in-the-loop）机制。&lt;/p&gt;

&lt;p&gt;上下文工程：数据语境化（将原始数据转化为模型可理解的上下文）、工具语境化（将工具接口封装为模型可调用的标准化接口）、上下文组织（按任务需求动态组装相关上下文）、Skills系统（可复用的能力模块）。&lt;/p&gt;

&lt;p&gt;交付质量保证：交付契约设计（明确任务的输入、输出、质量标准）、评估Rubrics（结构化的多维度评估框架）、自我检验（模型对自身输出进行质量审查）、交付过程透明化（用户可以查看智能体的推理过程和决策依据）。&lt;/p&gt;

&lt;p&gt;这套体系的目标是使能动模型可信、可靠地融入工作。&lt;/p&gt;

&lt;h3 id=&quot;六企业数字化的范式跃迁&quot;&gt;六、企业数字化的范式跃迁&lt;/h3&gt;
&lt;p&gt;从更长的时间尺度看，企业数字化经历了五个阶段的演进，每个阶段都可以用一个「System of X」的框架来概括。&lt;/p&gt;

&lt;p&gt;System of Information（信息系统阶段）：以微软Office为代表，企业开始了桌面数字化，核心是文档处理。&lt;/p&gt;

&lt;p&gt;System of Record（记录系统阶段）：ERP、CRM等企业级应用的普及，核心是记录业务流程和决策过程。&lt;/p&gt;

&lt;p&gt;System of Engagement（连接系统阶段）：互联网时代的企业数字化，核心是与用户、客户、渠道、供应商的在线连接，以及内部的邮件、协同工具。&lt;/p&gt;

&lt;p&gt;System of Insight（洞察系统阶段）：大数据和BI分析驱动的阶段，核心是从海量数据中提取洞察，支持数据驱动的决策。云原生、数据栈、低代码平台构成了这一阶段的技术基础。&lt;/p&gt;

&lt;p&gt;System of Intelligence（智能系统阶段）：AI时代的全新范式。多模态自动感知系统、从数据到模型的完整链路、辅助和全自动化的决策与执行能力，构成了这一阶段的核心特征。覆盖面从每个场景、每个职能、每个行业，延伸到每个设备、每个工艺、每个岗位。&lt;/p&gt;

&lt;p&gt;当前正在发生的范式转移最核心的特征可以用三个「从…到…」概括：从「数据驱动」到「语义驱动」，从「IT工具」到「业务语言」，从「辅助决策」到「智能决策」。软件3.0和上下文工程为企业提供了用统一的自然语言进行内部数字化整合的可能性。这一转变的意义堪比从纸质文档到电子文档的跃迁——它改变的不仅仅是效率，更是企业组织和运作的底层范式。&lt;/p&gt;

&lt;h3 id=&quot;七fde模式智能体落地的关键角色&quot;&gt;七、FDE模式：智能体落地的关键角色&lt;/h3&gt;
&lt;p&gt;FDE（Forward Deployed Engineer，前沿部署工程师）是一个源自Palantir的概念，正在成为AI时代智能体落地的核心角色。&lt;/p&gt;

&lt;p&gt;为什么需要FDE？ 根本原因在于「应用鸿沟」的存在。大语言模型具有强大的通用能力，但企业业务具有高度的特殊性——特定的数据格式、独特的业务流程、复杂的组织架构、微妙的行业知识。模型的通用性与业务的特殊性之间存在一道鸿沟，FDE的角色就是架设桥梁。&lt;/p&gt;

&lt;p&gt;在传统软件时代，这道鸿沟主要通过标准化产品+定制化实施来弥合。但AI时代的情况有所不同：需求高度不确定（业务方往往自己也不知道AI具体能帮他们做什么），迭代速度极快（模型能力每几个月就有显著提升），价值创造方式从交付固定功能转向持续交付可量化的业务价值。这些特征要求一种全新的角色——既懂技术又懂业务，既能做研究又能做工程，既能独立探索又能与客户共创。&lt;/p&gt;

&lt;p&gt;FDE的核心工作流是一个持续迭代的循环：需求理解（与领域专家沟通，深入理解业务场景）→ 业务转化（将业务需求转化为AI可处理的任务定义，包括提示词设计和工作流编排）→ 技术实施（设计输出格式、系统调优、API编排）→ 质量验证（测试、修正、达标确认）→ 生产部署（上线、监控、持续优化）。&lt;/p&gt;

&lt;p&gt;Palantir在空客的案例完美展示了FDE模式从0到1再到N的Scaling过程。2015年底，Palantir的FDE团队进驻空客工厂，通过深入一线访谈发现了数据孤岛问题，构建了SDDI数据集成方案，用Ontology（本体论）建立业务语义层，开发了第一个面向A350生产效率的MVP应用。2016年，FDE开始横向拓展，将相同的方法论复制到生产、供应链、调度、财务、质量等多个业务线。2017年，Palantir将这些经验抽象为Skywise平台产品（基础设施层Apollo+平台层Foundry+应用层），实现了规模化部署。2017年之后，平台开放给OEM、航空公司、供应商、MRO等整个航空生态系统。截至2024年，已有150+航空公司加入，平均每两周新增一家组织。&lt;/p&gt;

&lt;p&gt;这个案例的深层启示是：FDE模式的Scaling路径从「定制问题」起步，经由「抽象通用解」，最终「输出产品形态」。每一步都需要FDE在客户现场的深度浸入。&lt;/p&gt;

&lt;p&gt;OpenAI在John Deere（约翰迪尔）的案例则展示了FDE在农业领域的应用。John Deere的See &amp;amp; Spray智能喷洒技术利用36个摄像头和边缘算力识别杂草并进行精准喷洒，实现了化学农药使用减少60-70%、产量提升等显著成效。在这个项目中，OpenAI的FDE团队完成了从早期范围界定（系统调研、数据盘点、KPI定义）到验证（离线验证+边缘仿真+田间试验）再到交付（生产集成+监控运维+数据回流）的全流程。&lt;/p&gt;

&lt;h3 id=&quot;八fde在中国的落地挑战&quot;&gt;八、FDE在中国的落地挑战&lt;/h3&gt;
&lt;p&gt;FDE模式在中国的落地面临一系列独特的挑战。&lt;/p&gt;

&lt;p&gt;支付意愿问题。 中国企业对SaaS和知识服务的付费意愿长期偏低，习惯于为硬件和有形资产买单。FDE模式的核心价值是知识密集型的咨询+工程服务，这种价值形态在中国市场需要被重新定义和教育。&lt;/p&gt;

&lt;p&gt;数据基础设施碎片化。 中国企业的数字化历程以自研和定制化为主，形成了大量的「烟囱式」系统，数据孤岛问题严重。这直接导致FDE在企业端的集成工作从美国通常的2周延长到2-3个月，时间成本和人力成本大幅上升，解决方案的可复用性显著下降。&lt;/p&gt;

&lt;p&gt;文化差异。 美国的Palantir文化强调跨界、模糊边界的角色定义，FDE可以同时扮演工程师、顾问、研究员多重角色。中国企业文化倾向于清晰的岗位边界和职责划分，FDE的复合型角色定位在组织中可能面临定位困难。&lt;/p&gt;

&lt;p&gt;人才供给稀缺。 FDE要求极高的复合能力（业务理解+技术能力+产品思维），市场上这类人才极度稀缺。AI创业公司急需建立FDE团队以落地模型能力，传统企业服务商也需要FDE能力来支撑转型，供需之间的比例严重失衡。&lt;/p&gt;

&lt;p&gt;模型能力差异。 美国FDE可以直接调用前沿模型（Claude、GPT-5级别）的强大能力，重点精力放在业务集成和信任建立上。中国FDE往往需要在模型能力不足的情况下进行额外的补偿工作（微调、工程trick等），这进一步增加了项目的复杂度和成本。&lt;/p&gt;

&lt;h3 id=&quot;九fde的进化方向fdr与opc&quot;&gt;九、FDE的进化方向：FDR与OPC&lt;/h3&gt;
&lt;p&gt;FDE的角色定义本身也在快速进化。&lt;/p&gt;

&lt;p&gt;从FDE到FDR（Forward Deployed Researcher，前沿部署研究员）。 OpenAI和Anthropic正在推动的这一转变意味着：部署在客户现场的不仅仅是工程师，还是研究员。FDR的核心使命从「配置和集成现有技术」升级为「在客户环境中进行前沿技术的研究和创新」。在FDR模式下，客户项目本身成为了研究课题的来源，成功的技术模式会被快速产品化并反哺到平台的核心能力中。这种「双向知识流动」机制使得前沿研究与产业落地之间形成了更紧密的正反馈循环。&lt;/p&gt;

&lt;p&gt;一个典型的案例是企业知识库AI助手的开发：客户需求是「快速查询历史项目文档」，FDR在解决这一需求的过程中，探索了如何让LLM理解行业术语和概念，开发了新的「上下文压缩」技术和「多模态文档理解」方法，这些创新最终被整合到OpenAI的Assistants API中。&lt;/p&gt;

&lt;p&gt;从FDE到OPC（One Person Company，一人公司）。 这是一个更激进的进化方向。FDE和OPC创业者在能力要求上存在高度的重叠：全栈技术能力、深度业务理解、在模糊环境中定义问题的能力、跨界沟通协作、结果导向、快速适应。当AI工具足够强大时，一个具备FDE能力的个体完全可以独立创建和运营一家公司。&lt;/p&gt;

&lt;p&gt;Major Shlomo的Base44案例是这一模式最引人注目的证明。作为一人公司（直到被收购前最后一个月才招了第一名员工），他用AI构建MVP（2-4周完成一个产品），直接推送生产环境，通过内容营销和病毒传播机制实现增长，最终在6个月内被Wix以8000万美元现金收购。年度经常性收入350万美元，25万+用户，单月利润18.9万美元。&lt;/p&gt;

&lt;p&gt;这个案例表明：AI工具（尤其是Claude Code等强能动性工具）正在将FDE的能力「民主化」——将原本只有大型组织才能调动的复杂工程能力赋予个体创业者。&lt;/p&gt;

&lt;h3 id=&quot;十fde赋能产业转型的系统路径&quot;&gt;十、FDE赋能产业转型的系统路径&lt;/h3&gt;
&lt;p&gt;将FDE模式嵌入产业升级的整体框架中，可以描绘出一条「以点到面」的发展路径。&lt;/p&gt;

&lt;p&gt;在To-G（政府服务）方向，FDE可以帮助建立集中的智能化服务体系，系统性覆盖和提升政务场景。参考已有的政务案例（企业画像、惠企赋能、政策补贴智能匹配等），通过「智能体+一张图+FDE」的模式系统性赋能政府部门。&lt;/p&gt;

&lt;p&gt;在To-B（企业服务）方向，FDE可以扮演「新咨询」和「新集成商」的双重角色。新咨询意味着辅助企业完成AI时代的认知升级和战略转型；新集成商意味着以可复制的智能体方案带动产业整体转型。这种模式可以被类比为「中国新生产力的麦肯锡」——用AI原生的方法论帮助企业理解和应用前沿技术。&lt;/p&gt;

&lt;p&gt;覆盖的产业领域与「十五五」和「人工智能+」政策高度关联：集成电路、人工智能、生物医药、高端装备、新能源汽车等五大先导产业，以及六大重点行业。&lt;/p&gt;

&lt;h2 id=&quot;结语面向智能文明的新长征&quot;&gt;结语：面向智能文明的新长征&lt;/h2&gt;
&lt;p&gt;回到文章开头提出的命题：我们正站在能动性边际成本趋零的历史拐点上。这个拐点的意义远超技术本身。&lt;/p&gt;

&lt;p&gt;从个体层面看，当每个人都可以借助智能体获得复杂任务的执行能力时，人类个体的生产力上限将被根本性地提升。从组织层面看，企业的竞争优势将从「拥有多少人」转向「拥有多好的上下文和驾驶舱」。从国家层面看，AI战略的竞争不仅是技术和资本的竞争，更是组织范式和文明范式的竞争。从文明层面看，科学第四范式的全面到来意味着人类认知和创新的速度将进入一个全新的量级。&lt;/p&gt;

&lt;p&gt;历史不会简单地重复，但它的韵律往往惊人地相似。正如蒸汽机催生了工业文明、电力催生了电气文明、计算机催生了数字文明，通用智能正在催生一个全新的文明形态。在这个文明中，知识、认知和能动性将像电力和互联网一样成为无处不在的公共基础设施。&lt;/p&gt;

&lt;p&gt;中国的机遇在于：在这场文明级的转型中，抓住从基础设施建设到产业落地、从模型研发到人才培养的每一个关键节点，构建起具有自主可控能力的全栈AI体系。FDE模式、软件3.0、上下文工程、驾驶舱体系——这些概念不仅仅是技术框架，它们是新文明基础设施的组成部分。&lt;/p&gt;

&lt;p&gt;这是一场面向智能文明的新长征。征途漫漫，但方向已经清晰。&lt;/p&gt;

</content>
  </entry>
  
  <entry>
    <title>Where Do We Go in the Age of General Intelligence?</title>
    <link href="https://mochiaochen.github.io/en/writing/2026/07/where-do-we-go-in-the-age-of-general-intelligence/" rel="alternate" type="text/html"/>
    <published>2026-07-22T16:30:00+08:00</published>
    <updated>2026-07-22T16:30:00+08:00</updated>
    <id>https://mochiaochen.github.io/en/writing/2026/07/general-intelligence-age-en</id>
    <content type="html" xml:base="https://mochiaochen.github.io/en/writing/2026/07/where-do-we-go-in-the-age-of-general-intelligence/">&lt;h2 id=&quot;introduction-what-kind-of-historical-turning-point-are-we-standing-at&quot;&gt;Introduction: What kind of historical turning point are we standing at?&lt;/h2&gt;

&lt;p&gt;In 2025, artificial intelligence reached an unprecedented threshold. If the explosion of ChatGPT in 2023 made the world aware of the knowledge-emergence capabilities of large language models, the central story of 2025 shifted to a deeper proposition: &lt;strong&gt;the marginal cost of agency is approaching zero&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;The history of information technology reveals three structural inflection points at which marginal costs collapsed. Around 1995, the spread of the internet drove the marginal cost of obtaining information sharply downward. In 2023, large language models led by the GPT family pushed the cost of accessing public knowledge close to zero: anyone could obtain expert-level answers through natural-language conversation. In 2025, as the agent ecosystem began to take shape, we started to witness a third turning point—the ability to act itself becoming available at scale.&lt;/p&gt;

&lt;p&gt;Here, “agency” means the capacity of a system to understand intent autonomously, make plans, use tools, execute tasks, and adjust in response to feedback. As the cost of deploying this capacity continues to fall, the productive structure of human society will be fundamentally reshaped.&lt;/p&gt;

&lt;p&gt;Understanding the full scope of this transformation requires a full-stack analytical framework. This essay proceeds along three dimensions: the strategic landscape, development trends, and R&amp;amp;D models. They are not parallel topics, but a progression from macro to micro and from “why” to “how.”&lt;/p&gt;

&lt;h2 id=&quot;part-i-the-strategic-landscapethe-civilisational-contest-in-the-age-of-general-intelligence&quot;&gt;Part I: The strategic landscape—the civilisational contest in the age of general intelligence&lt;/h2&gt;

&lt;h3 id=&quot;1-the-full-stack-a-systemic-map-of-the-ai-industry&quot;&gt;1. The full stack: a systemic map of the AI industry&lt;/h3&gt;

&lt;p&gt;Understanding the strategic landscape begins with a full-stack map.&lt;/p&gt;

&lt;p&gt;At the bottom, the &lt;strong&gt;infrastructure layer&lt;/strong&gt; encompasses energy consumption, mineral resources, materials science, and equipment manufacturing. AI’s deepest dependencies are physical. Training a frontier model consumes electricity measured in gigawatt-hours, while continuous inference demands an equally formidable energy supply. The location and construction of compute infrastructure are therefore constrained by energy and geography. The logic of America’s Stargate project is to bind compute and energy infrastructure into a single complex. China, meanwhile, is seeking differentiated advantages through its global lead in photovoltaics and continued progress in nuclear technology, including small modular reactors and research into controlled fusion.&lt;/p&gt;

&lt;p&gt;The &lt;strong&gt;intelligent-compute layer&lt;/strong&gt; includes AI computing centres, intelligent clouds, and data centres. Competition is moving beyond raw scale toward the efficiency of compute-model co-design. Whereas traditional cloud computing centred on general-purpose CPUs, the AI era requires floating-point matrix operations on GPU tensor cores and is evolving toward heterogeneous multi-chip architectures.&lt;/p&gt;

&lt;p&gt;The &lt;strong&gt;model-training layer&lt;/strong&gt; is the most fiercely contested. It comprises cognitive, contextual, embodied, spatial, and scientific intelligence—not isolated directions, but a spectrum of capabilities that together constitute general intelligence.&lt;/p&gt;

&lt;p&gt;The &lt;strong&gt;intelligence-extension layer&lt;/strong&gt; includes agents, Software 3.0, applications, and devices: the interfaces through which AI capabilities diffuse into the real world. Above it, the &lt;strong&gt;industrial-structure&lt;/strong&gt; and &lt;strong&gt;diffusion&lt;/strong&gt; layers concern how AI reshapes markets and spreads across consumer, business, and government settings.&lt;/p&gt;

&lt;p&gt;This map matters because AI competition is never confined to one technological dimension. The strength of a country or region depends on its accumulated capabilities at every layer and on how efficiently those layers work together.&lt;/p&gt;

&lt;h3 id=&quot;2-new-sources-of-innovation-who-drives-the-frontier&quot;&gt;2. New sources of innovation: who drives the frontier?&lt;/h3&gt;

&lt;p&gt;At the top of the stack, the organisational form of the leading force deserves special attention.&lt;/p&gt;

&lt;p&gt;The United States and China display sharply different structures. Frontier development in the United States is driven largely by research-intensive, closed-model start-ups. OpenAI, Anthropic, xAI, Thinking Machines Lab, and SSI form a distinctive cluster combining the depth of academic research with the speed of product engineering. Typically founded or led by elite researchers, these companies tightly couple basic research with product development and sustain exceptional “research density” through enormous financing. NVIDIA and Google participate as infrastructure and platform leaders.&lt;/p&gt;

&lt;p&gt;China’s landscape is different. Major technology companies—ByteDance through Doubao, Tencent, and Alibaba—lead model development and deployment, while an open-source ecosystem led by DeepSeek provides an important complementary force. DeepSeek R1 marked a notable advance in reasoning, yet purely research-driven start-ups remain relatively scarce, reflecting deeper differences in talent systems, capital markets, and organisational culture.&lt;/p&gt;

&lt;p&gt;The &lt;strong&gt;researcher-founder&lt;/strong&gt; is therefore crucial. Under the fourth paradigm of science—computation-driven discovery—the path from basic research to industrial application has shortened dramatically. The conventional “industry–university–research” model assumed a linear transfer from university research to corporate application. AI research companies break that assumption: one organisation performs basic research, engineering implementation, and value creation simultaneously. OpenAI’s transformation from a non-profit laboratory into one of the world’s most highly valued AI companies is the most dramatic example.&lt;/p&gt;

&lt;p&gt;At least four conditions sustain this model: a deep talent system, mature capital markets, compute at scale, and a closed loop from research to product. OpenAI and Anthropic lead because they possess all four.&lt;/p&gt;

&lt;h3 id=&quot;3-a-new-globalisation-the-geopolitics-of-ai-diffusion&quot;&gt;3. A new globalisation: the geopolitics of AI diffusion&lt;/h3&gt;

&lt;p&gt;AI is rewriting the foundations of globalisation. America’s diffusion strategy exports its entire technology stack—from chip architecture and model systems to application ecosystems—and has already captured a dominant share of the global stack. Its strategic aim is long-term ecosystem control through technical standards. The Genesis Mission extends this logic by using AI to accelerate research in strategic fields such as nuclear energy, biotechnology, and advanced materials.&lt;/p&gt;

&lt;p&gt;China’s response appears in its “AI+” industrial policy and the idea of a “new Digital Silk Road.” AI has been elevated further in the Fifteenth Five-Year Plan, while a digital version of the Belt and Road seeks adoption of Chinese AI stacks across the Global South. At heart, this is a contest between two paths for diffusing technological civilisation.&lt;/p&gt;

&lt;p&gt;The underlying variable is changing, however. Traditional technological dominance relied on control of proprietary closed systems; open source is weakening that foundation. When DeepSeek released a high-performance reasoning model openly, it created both a technical and geopolitical shock: it showed that closed barriers could be bypassed and gave developers worldwide an alternative to the American stack.&lt;/p&gt;

&lt;h3 id=&quot;4-new-infrastructure-the-80-year-cycle-and-bubble-risk&quot;&gt;4. New infrastructure: the 80-year cycle and bubble risk&lt;/h3&gt;

&lt;p&gt;Historically, the infrastructure phase of a major technological revolution brings massive capital expenditure, a cyclical bubble, and a “new dawn” after the bubble bursts. Railways, electricity, and the internet all followed this pattern.&lt;/p&gt;

&lt;p&gt;The current AI build-out looks similar. Stargate represents an infrastructure-economy strategy: secure the commanding heights of the next decade through investment in compute and energy on a vast scale. Such concentration also carries bubble risk, especially when near-term application revenue cannot match infrastructure spending.&lt;/p&gt;

&lt;p&gt;Carlota Perez’s &lt;em&gt;Technological Revolutions and Financial Capital&lt;/em&gt; describes four phases: eruption, frenzy, synergy, and maturity. The frenzy phase features an excessive influx of financial capital, inflated asset prices, and bubbles. A crash does not mean that the technology has failed; it is more like a forced reallocation of resources that prepares the synergy phase. Seen through this lens, the current AI boom may be moving from eruption into frenzy.&lt;/p&gt;

&lt;p&gt;China offers a contrast: government-guided, more balanced construction of computing and training centres aims to avoid purely market-driven overinvestment. Its strength in photovoltaics and continued nuclear research provide a longer-term answer to AI’s energy demand. If the rough 80-year rhythm of paradigm-level revolutions still holds, we are at the beginning of a new cycle.&lt;/p&gt;

&lt;h3 id=&quot;5-new-digitalisation-platform-ecosystems-and-inflection-effects&quot;&gt;5. New digitalisation: platform ecosystems and inflection effects&lt;/h3&gt;

&lt;p&gt;The digital economy is driven by platforms and ecosystems. Every change in computing paradigm produces a new platform and an ecosystem around it: Microsoft and Intel in the PC era, Google and Amazon on the internet, Apple and WeChat in the mobile era.&lt;/p&gt;

&lt;p&gt;New AI platforms are now forming. The United States is building a global ecosystem on its lead in frontier models and cloud infrastructure. China has a distinct consumer advantage: its huge mobile-internet user base offers fertile ground for rapid AI diffusion. In enterprise software, China may have an unusual leapfrogging opportunity. Because many companies remain behind their American counterparts in conventional digitalisation, they may skip parts of the traditional SaaS stage and move directly to AI-native enterprise services.&lt;/p&gt;

&lt;h3 id=&quot;6-new-value-and-new-industries&quot;&gt;6. New value and new industries&lt;/h3&gt;

&lt;p&gt;AI’s industrial impact goes far beyond efficiency. China’s Fifteenth Five-Year Plan treats semiconductors, AI, intelligent manufacturing, quantum computing, and the low-altitude economy as strategic emerging industries. These sectors are tightly coupled: AI chips depend on semiconductor advances; intelligent manufacturing depends on embodied intelligence; the low-altitude economy is a natural arena for spatial intelligence.&lt;/p&gt;

&lt;p&gt;Another often-overlooked dimension is America’s need to reindustrialise. After decades of manufacturing hollowing-out, AI—especially embodied intelligence and industrial robotics—is viewed as a critical lever, and that need is shaping US policy priorities.&lt;/p&gt;

&lt;h3 id=&quot;7-civilisational-systems-a-change-in-the-paradigm-of-science&quot;&gt;7. Civilisational systems: a change in the paradigm of science&lt;/h3&gt;

&lt;p&gt;At civilisational scale, AI is driving a fourth scientific paradigm.&lt;/p&gt;

&lt;p&gt;The &lt;strong&gt;first paradigm&lt;/strong&gt; was empirical: observation and thought accumulated knowledge slowly through oral transmission, from Egyptian irrigation rules to China’s twenty-four solar terms. The &lt;strong&gt;second&lt;/strong&gt;, marked by Galileo’s systematic experiments, introduced controlled variables and reproducibility. The &lt;strong&gt;third&lt;/strong&gt; was theoretical, reaching its heights in Newtonian mechanics, Maxwell’s equations, and relativity; mathematical modelling and peer review became standard.&lt;/p&gt;

&lt;p&gt;The &lt;strong&gt;fourth paradigm&lt;/strong&gt;, proposed by Jim Gray in 2007, is computational. It combines algorithms, large-scale data collection, human-machine collaboration, and iterative verification. AlphaFold’s protein-structure predictions, end-to-end autonomous driving, and neural decoding in brain-computer interfaces are representative. Discovery can shrink from years and months to weeks and days, while disciplinary boundaries expand.&lt;/p&gt;

&lt;p&gt;This changes the organisation of innovation. Universities dominated the third paradigm, with research entering industry through technology-transfer mechanisms such as the Bayh–Dole Act. In the fourth, research-intensive start-ups become central because they combine research, engineering, and markets to accelerate the whole journey from minus one to one. OpenAI, DeepMind, and DeepSeek embody the model.&lt;/p&gt;

&lt;p&gt;Vannevar Bush’s 1944 report &lt;em&gt;Science, the Endless Frontier&lt;/em&gt; laid the strategic foundation of America’s post-war research system. Nearly eighty years later, AI is driving another reconstruction. America’s Genesis Mission targets advanced manufacturing, biotechnology, modern nuclear energy, fusion, and the electrical grid. China has elevated AI to the highest strategic level and is exploring a new research system. The new industry–university–research combination is, in essence, a rearrangement of innovative resources during the transition from the third paradigm to the fourth.&lt;/p&gt;

&lt;h3 id=&quot;8-technology-production-relations-environment-and-population&quot;&gt;8. Technology, production relations, environment, and population&lt;/h3&gt;

&lt;p&gt;AI is accelerating four technological frontiers: new energy, life sciences, materials, and space. All rely heavily on simulation, data, and AI-assisted experimental design. In materials science, for example, AI can screen millions of possible compounds for target properties and compress years of discovery into weeks.&lt;/p&gt;

&lt;p&gt;The financial system is becoming a key production variable. A developed chain from venture capital to IPO exits supports research-intensive companies. At the same time, structural pressure on the dollar’s reserve status, the internationalisation of the renminbi, and digital currencies including stablecoins are changing global capital flows. AI is also transforming finance itself, from quantitative trading and risk assessment to compliance monitoring.&lt;/p&gt;

&lt;p&gt;The expansion of human activity from the ground to low altitude, high altitude, and space is creating new industrial ecosystems. Drone logistics, urban air mobility, and airspace management have become Chinese policy priorities. Integrated aerospace transport and information systems are moving from science fiction toward engineering reality. Their enabling technology is spatial intelligence: AI’s ability to understand, reason about, and plan in three-dimensional space.&lt;/p&gt;

&lt;p&gt;China also faces a deep demographic challenge: more than 310 million people are over sixty and the birth rate has fallen to 6.77 per thousand. This creates both demand and constraint. Education requires AI-driven reform for new forms of talent; healthcare faces enormous demand in imaging, remote surgery, and intelligent elder care. In this sense, AI is a strategic tool for responding to ageing.&lt;/p&gt;

&lt;h2 id=&quot;part-ii-development-trendsfrom-general-intelligence-to-agents-and-new-digitalisation&quot;&gt;Part II: Development trends—from general intelligence to agents and new digitalisation&lt;/h2&gt;

&lt;h3 id=&quot;1-five-stages-from-learning-knowledge-to-learning-organisation&quot;&gt;1. Five stages: from learning knowledge to learning organisation&lt;/h3&gt;

&lt;p&gt;The development of general intelligence can be divided into five qualitative stages.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Stage one: learning knowledge.&lt;/strong&gt; Pre-training absorbs public knowledge from vast internet corpora. GPT-3 demonstrated surprising knowledge emergence, and Doubao also performed strongly at this stage. ChatGPT’s explosive adoption marked its maturity: for the first time, people could access knowledge from almost any field through natural-language dialogue. Scaling here centred on context length and parameter count.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Stage two: learning to reason.&lt;/strong&gt; Since 2024, post-training has sharply improved reasoning. OpenAI’s o-series and DeepSeek R1 showed how reinforcement learning and self-reflection can strengthen logic and problem solving. DeepSeek’s open-source strategy accelerated the global diffusion of reasoning, while products such as Deep Research began to create practical value. Scaling shifted toward the iterative depth of self-consistent reflection.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Stage three: learning to act.&lt;/strong&gt; This was the central battlefield of 2025. Continual training must combine intent understanding, instruction following, long-horizon planning, and tool use. Agency has two branches: digital-world agency, represented by Claude Code across product design, research, development, marketing, and operations; and physical-world agency in robots, autonomous vehicles, and other embodied systems. Anthropic’s dramatic valuation growth reflected the market’s pricing of agency. Claude Code showed that value grows non-linearly once a model can complete complex work autonomously inside a real development environment. In the physical world, multiple paths toward world models are beginning to provide embodied systems with physical intuition. Scaling becomes a trinity of environment, tools, and education.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Stage four: learning to innovate.&lt;/strong&gt; Still early, this stage seeks genuine invention beyond existing knowledge. AI-assisted AI development and AI Scientists have entered experimentation, while mathematical reasoning at IMO gold-medal level hints at the potential. Scaling will depend on accumulated time spent innovating.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Stage five: learning to organise.&lt;/strong&gt; The ultimate form is a multi-agent system capable of complex decisions, organisational coordination, and strategy. Research on multi-agent coordination remains nascent, but its long-term importance may equal that of single-agent breakthroughs. When highly autonomous agents can divide labour like a human organisation, a new social form will emerge. Scaling will depend on the complexity of adaptation to the environment.&lt;/p&gt;

&lt;h3 id=&quot;2-a-new-path-for-scaling-laws-from-models-to-agents&quot;&gt;2. A new path for scaling laws: from models to agents&lt;/h3&gt;

&lt;p&gt;Scaling laws have expanded profoundly. Early work by Kaplan et al. (2020) and Hoffmann et al. (2022) focused on power-law relationships among parameters, data, and compute in pre-training. As capabilities and use cases deepen, new scaling dimensions are appearing.&lt;/p&gt;

&lt;p&gt;One useful analogy compares AI with the social development of civilisation: genetic inheritance, language, education, industrialisation, digitalisation, and finally intelligence each improved the transmission and amplification of capability. AI follows a similar path—from data-driven “inheritance” in pre-training, to reasoning-driven “cognition” in post-training, to interaction-driven “agency” in continual training, and ultimately to higher-order scaling in innovation and organisation.&lt;/p&gt;

&lt;p&gt;The speed and path of this new scaling depend heavily on the maturity of digital environments and infrastructure. Competition in the agent era is therefore not only about the model, but about the quality and completeness of its environment: tools, data, cockpit, and context engineering.&lt;/p&gt;

&lt;h3 id=&quot;3-the-technical-frontier-of-general-intelligence&quot;&gt;3. The technical frontier of general intelligence&lt;/h3&gt;

&lt;p&gt;At the model foundation, sparse attention is improving efficiency and enabling extremely long contexts; unified architectures seek high-quality understanding and multimodal generation in one model; diffusion and flow-matching methods are advancing image, video, and audio generation; long-context and memory breakthroughs support million-token inputs; evaluation is moving from standard benchmarks to real-world, agency-oriented tests; and continual learning seeks new capabilities without catastrophic forgetting.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Cognitive intelligence&lt;/strong&gt; pursues difficult problem solving and deep reflection. Its frontiers include AI-assisted AI development, end-to-end AI Scientists that propose hypotheses and design and analyse experiments, and a deeper union of cognition and agency so that models can both think and act. Superintelligence remains distant, but leading laboratories are already researching its safety and alignment.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Contextual intelligence&lt;/strong&gt; concerns perception and understanding of real situations. Omnimodal integration across audio, video, touch, and other channels is a central challenge. “Deep Context” means understanding not only language but the full semantic structure of tasks and rewards. As agents operate in increasingly complex environments, demand for native contextual data is exploding.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Embodied intelligence&lt;/strong&gt; has accelerated. Humanoid robots such as Figure and Unitree, quadrupeds for education, entertainment, and inspection, industrial robots, and autonomous driving form four major settings. Architectures are moving from modular pipelines to end-to-end systems, with world models and vision-language-action models as key directions. Work such as Physical Intelligence’s π0.5 suggests that language-model reasoning can transfer into physical manipulation. Data—embodiment, process, and annotation—remains the central bottleneck to scale.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Spatial intelligence&lt;/strong&gt; is rapidly emerging. In digital space, systems such as World Labs and Marble build three-dimensional world models; in physical space, companies such as Qidaisong focus on contextual understanding and spatial-intelligence machines. Better spatial reasoning and attention will directly accelerate embodied deployment.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Scientific intelligence&lt;/strong&gt;, or AI for Science, is among the most strategically valuable applications. It requires dedicated scientific data and model systems, agents participating end to end in research, and AI Scientists capable of proposing and testing hypotheses. The Genesis Mission represents national-level US investment. The ultimate aim is to accelerate every step of discovery, from data collection and hypothesis generation to experimental validation.&lt;/p&gt;

&lt;h3 id=&quot;4-agents-the-full-release-of-agency&quot;&gt;4. Agents: the full release of agency&lt;/h3&gt;

&lt;p&gt;2025 was widely called “the year of the agent.” Four capabilities define an agent: understanding the user’s real intent, following instructions faithfully, decomposing complex goals into executable long-range plans, and interacting with tools such as APIs, software interfaces, and file systems.&lt;/p&gt;

&lt;p&gt;Releasing this agency depends on three conditions.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Real environments.&lt;/strong&gt; Agents must learn through practice in real settings. The shift from static internet text to the complete context of valuable real-world situations is essential. Agents need to complete long, multi-turn tasks independently and improve through interaction.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Real evaluation.&lt;/strong&gt; Static benchmarks no longer measure real capability effectively. Inputs and outputs must come from live data and authentic user needs; evaluations must change dynamically and introduce recent queries to prevent reward hacking and test-set overfitting.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Real data.&lt;/strong&gt; Data collection is reorganising around environmental context. The decisive move from generic internet context to complete, multidimensional, high-value contexts is what turns an agent from merely usable into genuinely good.&lt;/p&gt;

&lt;h3 id=&quot;5-the-cockpit-how-humans-control-agency&quot;&gt;5. The cockpit: how humans control agency&lt;/h3&gt;

&lt;p&gt;In Lu Qi’s framework, the &lt;strong&gt;cockpit&lt;/strong&gt; is the interface through which people use and control an agent’s agency.&lt;/p&gt;

&lt;p&gt;History clarifies the idea. The internal-combustion engine gave humanity powerful physical agency, but releasing it required a long evolution of cockpits: horse, carriage, Ford Model T, and highway system. Each iteration expanded the range and efficiency of mechanical power. The internet likewise provided informational agency, while its cockpit evolved from servers to browsers, search engines, and social platforms.&lt;/p&gt;

&lt;p&gt;AI follows the same law. Models possess enormous latent agency, but cockpit iteration and engineering systems are needed to release it. OpenAI Codex was a first-generation programming cockpit; Claude Code marked a major engineering advance; the ecosystem forming around it signals further maturity.&lt;/p&gt;

&lt;p&gt;Claude Code’s core innovation is a shared context between silicon and carbon. Earlier cloud-first tools asked developers to move their environment to the cloud, yet most real work remained local—in files, terminals, company databases, and internal resources. Claude Code instead installs on the developer’s machine and can work directly in that environment. Directory-level Markdown instructions describe local context so AI can see and use the same world as the human.&lt;/p&gt;

&lt;p&gt;The deeper principle is that AI and humans should share one working context. Rather than building an isolated digital environment for AI, this design embeds AI into the environment people already inhabit. It reduces integration complexity and makes human-agent collaboration natural.&lt;/p&gt;

&lt;p&gt;Cockpit evolution resembles levels of autonomous driving. In the assistance stage, humans lead and systems help; early conversational models paired with workflow frameworks such as Coze and LangChain belong here. In the conditional-autonomy stage, the system leads but humans can intervene; agentic reasoning models with long-horizon planning and environmental interaction are moving into this phase. The long-term destination is full autonomy: system-led, zero intervention, and high agency.&lt;/p&gt;

&lt;p&gt;Each stage has distinct engineering problems. Assistance requires workflow orchestration to compensate for weak multi-turn and long-horizon reasoning. Conditional autonomy requires safe execution—sandboxes, permission audits, behavioural guardrails, and humans in the loop—plus context engineering. Full autonomy requires mature quality assurance: delivery contracts, evaluation rubrics, self-checks, and transparent processes.&lt;/p&gt;

&lt;h3 id=&quot;6-context-engineering-the-operating-system-for-agents&quot;&gt;6. Context engineering: the operating system for agents&lt;/h3&gt;

&lt;p&gt;If the cockpit is the hardware interface between human and agent, &lt;strong&gt;context engineering&lt;/strong&gt; is the agent’s operating system. It has three core layers.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;System prompts&lt;/strong&gt; carry organisational purpose and personal context. They define who an agent is, whom it serves, and which principles it follows.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Structure prompts&lt;/strong&gt; carry static environmental context. Like a physical office, they define the workspace, the function of each area, and the responsibilities of different organisational roles.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Task prompts&lt;/strong&gt; carry dynamic context, from simple workflows to sophisticated cognitive scaffolding. Scaffolding does more than say what to do; it supplies a way to think—problem definition, construction of the solution space, decomposition into cognitive dimensions, and a systematic reasoning process.&lt;/p&gt;

&lt;p&gt;The complete system also includes the cockpit beneath it—operating system, files, and hardware—and a context-processing loop: perceive the environment, analyse and plan, execute, remember state, and improve through feedback.&lt;/p&gt;

&lt;h3 id=&quot;7-new-digitalisation-the-arrival-of-software-30&quot;&gt;7. New digitalisation: the arrival of Software 3.0&lt;/h3&gt;

&lt;p&gt;Software 1.0 was the symbolic era: code as input, compiler as executor. Software 2.0 was the data era: data-driven neural inference. Software 3.0 is the language era: natural-language description, understood and executed by large models.&lt;/p&gt;

&lt;p&gt;Its stack has five layers:&lt;/p&gt;

&lt;ol&gt;
  &lt;li&gt;&lt;strong&gt;Hardware compute:&lt;/strong&gt; floating-point matrix operations on GPU tensor cores, with resources allocated across inference tasks through time-sharing.&lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;System core:&lt;/strong&gt; language and world models form a new “context computer,” where prompts are input, inference is computation, and generated tokens are output.&lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Commercial OS:&lt;/strong&gt; cockpit plus context engineering. In the physical world, a spatial cockpit plus physical context forms a “spatial-intelligence machine.”&lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Development extensions:&lt;/strong&gt; plugins, skills, hooks, and the Model Context Protocol form an application ecosystem. MCP is becoming a standard way for agents to access tools and services.&lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Application ecosystem:&lt;/strong&gt; digital agents—agentic, customised, sometimes disposable apps—and embodied agents such as robots, autonomous vehicles, and new devices serve end users.&lt;/li&gt;
&lt;/ol&gt;

&lt;h3 id=&quot;8-reshaping-industrial-structure&quot;&gt;8. Reshaping industrial structure&lt;/h3&gt;

&lt;p&gt;Software 3.0 affects the entire industrial structure. In consumer markets, AI is rebuilding gateways across communication, social media, content, games, education, healthcare, finance, and housing. The change reaches beyond smarter interfaces to the underlying service logic: from standardised products to personalised intelligent services.&lt;/p&gt;

&lt;p&gt;In enterprise markets, new IaaS for domestic compute and training, new PaaS for model-compute coordination and context efficiency, and new SaaS for agentic applications form a three-tier architecture. Cars, robots, drones, autonomous ships, AI glasses, wearables, and phones become physical extensions of AI.&lt;/p&gt;

&lt;p&gt;At the macro level, an “AI factory” is taking shape. Supply chains, sales, customer support, employee experience, finance, legal work, administration, and government relations will all connect through AI cockpits to a unified intelligent operating system.&lt;/p&gt;

&lt;h2 id=&quot;part-iii-rd-modelsfrom-the-laboratory-to-the-battlefield&quot;&gt;Part III: R&amp;amp;D models—from the laboratory to the battlefield&lt;/h2&gt;

&lt;h3 id=&quot;1-six-elements-of-a-new-rd-model&quot;&gt;1. Six elements of a new R&amp;amp;D model&lt;/h3&gt;

&lt;p&gt;The new model has six pillars: a compute system built around the agent cloud; a model system with closed-loop agency; a data system based on context engineering; a cockpit that releases agency; a development and delivery model based on efficient, customised co-creation; and an application model based on systematic deployment.&lt;/p&gt;

&lt;h3 id=&quot;2-a-new-model-system-open-agentic-models-and-an-evaluation-ecosystem&quot;&gt;2. A new model system: open agentic models and an evaluation ecosystem&lt;/h3&gt;

&lt;p&gt;Several directions matter.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Open agentic models.&lt;/strong&gt; On open foundations such as Qwen3 and DeepSeek, developers must build intent understanding, goal adherence, planning and decomposition, and interactive reasoning. These capabilities should generalise across tool use, web interaction, deep research, and coding.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Real-world evaluation through an Open Agent Arena.&lt;/strong&gt; Static benchmarks are insufficient. Dynamic evaluations must draw on authentic user needs and refresh regularly to prevent overfitting.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Training as a Service.&lt;/strong&gt; High-value industrial scenarios should form a closed loop with model training: real settings supply training data and evaluation standards; trained models return to those settings for validation and continued iteration.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Next-generation agency systems.&lt;/strong&gt; Agent-development frameworks such as NexAU can support end-to-end agentic data synthesis and reinforcement-learning infrastructure such as NexRL and NexVenusCL, moving agentic models from experiments to industrial deployment.&lt;/p&gt;

&lt;h3 id=&quot;3-a-new-compute-system-from-compute-cloud-to-agent-cloud&quot;&gt;3. A new compute system: from compute cloud to agent cloud&lt;/h3&gt;

&lt;p&gt;The value logic of compute is changing. The old model rented GPU time slices, with thin margins, limited barriers, and severe commoditisation. “Token as a Service” raises value density by charging for inference but remains dependent on replaceable public internet context.&lt;/p&gt;

&lt;p&gt;The new core is the &lt;strong&gt;agent cloud&lt;/strong&gt;, a compute service organised around agent execution. Its proposition is “context as a service plus agent operation”: serve agents rather than people, move from internet context to valuable enterprise context, and replace usage-based billing with outcome-oriented pricing. Enterprise productivity context compounds—the longer a system is used, the richer its accumulated context and the more valuable the service.&lt;/p&gt;

&lt;p&gt;The architecture has three layers. New IaaS provides intelligent-cloud infrastructure, advanced and domestic chips, training compute, and sandboxes. New PaaS provides agent runtimes, context engineering as a service, and permission and security management. New SaaS provides enterprise-productivity services, agent engineering, workspaces, agency configuration, and agent cockpits.&lt;/p&gt;

&lt;p&gt;Across all three lies model-compute co-design: coordinating heterogeneous multi-chip systems with model architecture, plus Training as a Service under continual learning through automatic optimisation, reinforcement learning, and asynchronous parallelism.&lt;/p&gt;

&lt;h3 id=&quot;4-a-new-data-system-from-internet-data-to-environmental-interaction&quot;&gt;4. A new data system: from internet data to environmental interaction&lt;/h3&gt;

&lt;p&gt;Data requirements change with agency. Pre-training relied on internet text and books; post-training turned to conversations and reasoning traces; agency training needs context and feedback from environmental interaction. Such data comes from real work settings, business processes, and decisions, and is highly situation-specific and time-sensitive.&lt;/p&gt;

&lt;p&gt;American model and data companies are systematically building occupational data assets. Anthropic and Stanford use the US Department of Labor’s O*NET database to study AI’s impact on occupations and answer a strategic question: which jobs should AI learn first?&lt;/p&gt;

&lt;p&gt;Surge AI is recreating full enterprise software environments and organisational structures, generating tens of thousands of logically consistent synthetic records. These “synthetic enterprises” offer agents safe, controllable, scalable places to practise.&lt;/p&gt;

&lt;p&gt;OpenAI’s reported Mercury project has recruited Wall Street professionals at $150 an hour to annotate financial data, illustrating the scarcity and cost of high-value domain knowledge. Together these cases point to one conclusion: competition over agentic data is a contest to construct the most complete, authentic, and valuable situational context.&lt;/p&gt;

&lt;h3 id=&quot;5-engineering-the-cockpit&quot;&gt;5. Engineering the cockpit&lt;/h3&gt;

&lt;p&gt;The central design principle is to harness agency while compensating for its weaknesses: behavioural drift over long tasks, insufficient business understanding, and unstable delivery quality.&lt;/p&gt;

&lt;p&gt;The new cockpit builds four kinds of protection and enhancement:&lt;/p&gt;

&lt;ul&gt;
  &lt;li&gt;&lt;strong&gt;Safe execution:&lt;/strong&gt; sandboxes, permission auditing, behavioural guardrails, and human-in-the-loop controls.&lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Context engineering:&lt;/strong&gt; contextualising raw data and tools, dynamically organising relevant context, and packaging reusable capabilities as skills.&lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Delivery quality:&lt;/strong&gt; explicit input, output, and quality contracts; multidimensional evaluation rubrics; model self-checks; and transparent delivery processes.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The goal is to integrate agentic models into work in a trustworthy and reliable way.&lt;/p&gt;

&lt;h3 id=&quot;6-a-paradigm-shift-in-enterprise-digitalisation&quot;&gt;6. A paradigm shift in enterprise digitalisation&lt;/h3&gt;

&lt;p&gt;Enterprise digitalisation can be described in five stages.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;System of Information:&lt;/strong&gt; desktop digitalisation led by Microsoft Office, centred on documents.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;System of Record:&lt;/strong&gt; ERP and CRM record business processes and decisions.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;System of Engagement:&lt;/strong&gt; internet-era connections with customers, channels, suppliers, and employees through email and collaboration tools.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;System of Insight:&lt;/strong&gt; big data and BI extract insights for data-driven decisions, supported by cloud-native infrastructure, data stacks, and low-code platforms.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;System of Intelligence:&lt;/strong&gt; multimodal perception, complete data-to-model pipelines, and partially or fully automated decisions and execution across every function, industry, device, process, and role.&lt;/p&gt;

&lt;p&gt;The present shift can be summarised in three movements: from data-driven to semantics-driven; from IT tools to business language; and from decision support to intelligent decision-making. Software 3.0 and context engineering make it possible to integrate enterprise digitalisation through a shared natural language. Like the move from paper to electronic documents, this changes not merely efficiency but the foundations of organisation and operation.&lt;/p&gt;

&lt;h3 id=&quot;7-fde-the-key-role-in-agent-deployment&quot;&gt;7. FDE: the key role in agent deployment&lt;/h3&gt;

&lt;p&gt;The &lt;strong&gt;Forward Deployed Engineer&lt;/strong&gt;, a role pioneered by Palantir, is becoming central to the agent era.&lt;/p&gt;

&lt;p&gt;FDEs are needed because of the application gap. Large models are general; every business is specific, with its own data formats, processes, organisational structures, and tacit domain knowledge. In traditional software, standard products plus custom implementation bridged the gap. AI adds uncertain requirements, rapid model improvement, and a shift from fixed functionality to continuously measurable value. The new role must understand technology and business, combine research and engineering, explore independently, and co-create with customers.&lt;/p&gt;

&lt;p&gt;The workflow is iterative: understand needs with domain experts; translate them into AI task definitions, prompts, and workflows; implement output formats, tuning, and API orchestration; verify and correct quality; then deploy, monitor, and improve in production.&lt;/p&gt;

&lt;p&gt;Palantir’s work with Airbus shows the path from zero to one to many. In late 2015, FDEs entered Airbus factories, found data silos, built an SDDI integration solution, used an ontology to create a semantic business layer, and developed an MVP for A350 production efficiency. In 2016, the approach expanded across production, supply chains, scheduling, finance, and quality. In 2017, these lessons became the Skywise platform—Apollo infrastructure, Foundry platform, and applications—and then opened to OEMs, airlines, suppliers, and MRO providers. By 2024, more than 150 airlines had joined, with an organisation added roughly every two weeks.&lt;/p&gt;

&lt;p&gt;The deeper lesson is that FDE scaling begins with a custom problem, abstracts a general solution, and ultimately produces a product. Every step depends on immersion at the customer site.&lt;/p&gt;

&lt;p&gt;OpenAI’s work with John Deere offers an agricultural example. Deere’s See &amp;amp; Spray system uses 36 cameras and edge compute to identify weeds and spray precisely, reportedly reducing chemical use by 60–70 percent while improving outcomes. The FDE workflow spans initial scoping, system and data review, and KPI definition; offline validation, edge simulation, and field trials; and finally production integration, monitoring, operations, and data feedback.&lt;/p&gt;

&lt;h3 id=&quot;8-the-challenge-of-fde-in-china&quot;&gt;8. The challenge of FDE in China&lt;/h3&gt;

&lt;p&gt;China presents distinct obstacles.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Willingness to pay.&lt;/strong&gt; Chinese enterprises have traditionally paid less readily for SaaS and knowledge services than for hardware and tangible assets. The knowledge-intensive blend of consulting and engineering must be redefined and explained.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Fragmented data infrastructure.&lt;/strong&gt; Years of in-house and customised systems have produced silos. Integration that may take two weeks in the United States can stretch to two or three months, raising labour costs and reducing reuse.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Cultural differences.&lt;/strong&gt; Palantir-style FDEs cross the boundaries of engineer, consultant, and researcher. Chinese organisations often favour clearly separated responsibilities, making a hybrid role harder to place.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Scarce talent.&lt;/strong&gt; The role demands business understanding, technical depth, product judgement, and comfort with ambiguity. AI start-ups need such teams to deploy models, while traditional service companies need them for transformation; supply falls far short of demand.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Model capability gaps.&lt;/strong&gt; American FDEs can often concentrate on integration and trust while using frontier models. Chinese FDEs may need extra fine-tuning and engineering work to compensate for weaker base capabilities, increasing complexity and cost.&lt;/p&gt;

&lt;h3 id=&quot;9-from-fde-to-fdr-and-opc&quot;&gt;9. From FDE to FDR and OPC&lt;/h3&gt;

&lt;p&gt;The role itself is evolving.&lt;/p&gt;

&lt;p&gt;The &lt;strong&gt;Forward Deployed Researcher&lt;/strong&gt; moves beyond configuring existing technology. OpenAI and Anthropic are putting researchers in customer environments so that customer problems become research questions. Successful methods are productised and fed back into the platform, creating a two-way flow between frontier research and industrial deployment.&lt;/p&gt;

&lt;p&gt;Consider an enterprise knowledge assistant. A request to search historical project documents may lead an FDR to study how models understand industry terminology, develop new context-compression and multimodal-document techniques, and eventually integrate those innovations into a broader API platform.&lt;/p&gt;

&lt;p&gt;The more radical evolution is the &lt;strong&gt;One Person Company&lt;/strong&gt;. FDEs and solo founders share full-stack technical ability, deep business understanding, problem definition under ambiguity, cross-domain communication, results orientation, and adaptability. When AI tools become powerful enough, one person with FDE capabilities can build and run a company.&lt;/p&gt;

&lt;p&gt;The Base44 case is a striking illustration. Founder Maor Shlomo reportedly operated essentially alone until hiring a first employee shortly before acquisition, used AI to build MVPs in two to four weeks, deployed directly to production, and grew through content and viral distribution. Wix acquired the company six months later for $80 million in cash. Reported annual recurring revenue was $3.5 million, with more than 250,000 users and monthly profit of $189,000.&lt;/p&gt;

&lt;p&gt;The point is not merely the headline numbers. Agentic tools such as Claude Code are democratising the complex engineering capacity once available only to large organisations.&lt;/p&gt;

&lt;h3 id=&quot;10-a-systemic-path-for-fde-led-industrial-transformation&quot;&gt;10. A systemic path for FDE-led industrial transformation&lt;/h3&gt;

&lt;p&gt;Embedded in industrial upgrading, FDE can support a path from point solutions to systemic transformation.&lt;/p&gt;

&lt;p&gt;In government services, FDEs can help build centralised intelligent systems across public-service scenarios. Existing patterns—enterprise profiling, business support, and intelligent matching of policy subsidies—suggest a model combining agents, a unified operational map, and FDE delivery.&lt;/p&gt;

&lt;p&gt;In enterprise services, FDEs can act as both “new consultants” and “new integrators.” Consulting helps organisations update their understanding and strategy for the AI era; integration uses reusable agent solutions to drive wider transformation. The aspiration is a kind of AI-native McKinsey for China’s new productive forces.&lt;/p&gt;

&lt;p&gt;The relevant industries align closely with the Fifteenth Five-Year Plan and “AI+” policy: integrated circuits, artificial intelligence, biomedicine, advanced equipment, new-energy vehicles, and other priority sectors.&lt;/p&gt;

&lt;h2 id=&quot;conclusion-a-new-long-march-toward-intelligent-civilisation&quot;&gt;Conclusion: a new Long March toward intelligent civilisation&lt;/h2&gt;

&lt;p&gt;We return to the opening proposition: we stand at a historical turning point where the marginal cost of agency is approaching zero. Its meaning extends far beyond technology.&lt;/p&gt;

&lt;p&gt;For individuals, access to agents raises the upper bound of what one person can execute. For organisations, advantage shifts from how many people they employ to the quality of their context and cockpit. For countries, AI competition is a contest not only of technology and capital but of organisational and civilisational paradigms. For civilisation, the arrival of the fourth scientific paradigm moves the speed of cognition and innovation to a new order of magnitude.&lt;/p&gt;

&lt;p&gt;History does not repeat itself simply, but its rhythms are often strikingly similar. The steam engine gave rise to industrial civilisation, electricity to the electrical age, and computers to digital civilisation. General intelligence is giving rise to another form, in which knowledge, cognition, and agency become ubiquitous infrastructure like electricity and the internet.&lt;/p&gt;

&lt;p&gt;China’s opportunity is to seize every critical link in this civilisational transformation—from infrastructure to industrial deployment and from model research to talent development—and build an autonomous, controllable full-stack AI system. FDE, Software 3.0, context engineering, and cockpit systems are not merely technical frameworks; they are components of the infrastructure of a new civilisation.&lt;/p&gt;

&lt;p&gt;This is a new Long March toward intelligent civilisation. The road is long, but the direction is clear.&lt;/p&gt;
</content>
  </entry>
  
  <entry>
    <title>算力越低，认知越高？</title>
    <link href="https://mochiaochen.github.io/writing/2026/07/lower-compute-higher-cognition/" rel="alternate" type="text/html"/>
    <published>2026-07-22T15:30:00+08:00</published>
    <updated>2026-07-22T15:30:00+08:00</updated>
    <id>https://mochiaochen.github.io/writing/2026/07/lower-compute-higher-cognition</id>
    <content type="html" xml:base="https://mochiaochen.github.io/writing/2026/07/lower-compute-higher-cognition/">&lt;p&gt;假设宇宙中存在一个被称为「拉普拉斯妖」&lt;sup id=&quot;fnref:1&quot; role=&quot;doc-noteref&quot;&gt;&lt;a href=&quot;#fn:1&quot; class=&quot;footnote&quot; rel=&quot;footnote&quot;&gt;1&lt;/a&gt;&lt;/sup&gt;的幽灵。它拥有无限的计算能力，掌握了宇宙中所有原子的动量和位置。你可能会觉得，这个生灵一定拥有宇宙中至高无上的智慧，能看透。&lt;/p&gt;

&lt;p&gt;真实情况完全出乎我们的意料。数学和计算机科学告诉我们一个冷酷的事实：&lt;mark&gt;拥有无限算力，意味着它根本不需要具备人类所定义的「智慧」。&lt;/mark&gt;在拉普拉斯妖的眼中，这个世界没有「苹果」，没有「国家」，没有「爱情」，只有一堆按照物理定律枯燥运动的原子。它不需要总结规律，因为它能直接暴力推演一切。&lt;/p&gt;

&lt;p&gt;近期，arXiv 上出现了一篇极具思想冲击力的论文，题为 &lt;em&gt;From Entropy to Epiplexity&lt;/em&gt;。几位来自卡内基梅隆大学和纽约大学的研究者重新审视了经典信息论，并提出了一个极度反常识的结论：&lt;strong&gt;如果你想在这个世界上学到真正有用的知识体系，建立起所谓的高级认知，你必须是一个「算力受限」的观察者。&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;这篇论文虽然通篇在用严格的数学公式和密码学概念讨论大语言模型（LLM）和机器学习，它实际上揭示了一套极高深的人生算法。&lt;/p&gt;

&lt;h2 id=&quot;一shannon-的死穴与雪花屏的诅咒&quot;&gt;一、Shannon 的死穴与雪花屏的诅咒&lt;/h2&gt;

&lt;p&gt;在经典的信息论中，信息的单位是比特。Claude Shannon&lt;sup id=&quot;fnref:2&quot; role=&quot;doc-noteref&quot;&gt;&lt;a href=&quot;#fn:2&quot; class=&quot;footnote&quot; rel=&quot;footnote&quot;&gt;2&lt;/a&gt;&lt;/sup&gt; 告诉我们，信息的本质是「消除不确定性」。一个事件越是不可预测、越显得混乱，它蕴含的信息量就越大。这就是所谓的「信息熵」（Entropy）。在这个标准下，我们每天都在被海量的信息无情轰炸。&lt;/p&gt;

&lt;p&gt;或许有人会说，我们要获取更多的信息才能找到一个「正确」的人生路径。但想象一台没有插天线的旧电视机，屏幕上全是密密麻麻的雪花噪点。它的画面是完全随机的，你永远无法预测下一个像素是黑还是白。根据 Shannon 的理论，以及后来苏联数学家 Kolmogorov 的复杂性理论，这台雪花屏包含的信息量极大，几乎达到了理论上的极值。&lt;/p&gt;

&lt;blockquote&gt;
  &lt;p&gt;但绝对没有任何正常人会盯着雪花屏看上一整天。因为那里面毫无「意义」可言。&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;为了填补这个漏洞，论文作者引入了两个全新的核心概念：&lt;strong&gt;Time-bounded Entropy（时间受限熵）&lt;/strong&gt;和 &lt;strong&gt;Epiplexity&lt;/strong&gt;。&lt;/p&gt;

&lt;p&gt;时间受限熵，指的是在有限的时间和算力下，你觉得完全随机、毫无头绪的不可预测内容。生活中的绝大多数琐事、社交媒体上的情绪宣泄、股市里每天高频的随机波动，对普通人来说都属于这种极高的时间受限熵。你投入再多的精力去追踪，也提炼不出任何可复用的规律。你只是在消耗生命，去计算一组类似于密码学中的伪随机数。&lt;/p&gt;

&lt;p&gt;与之相对的，Epiplexity 这个词可以精准地翻译为「结构复杂性」&lt;sup id=&quot;fnref:3&quot; role=&quot;doc-noteref&quot;&gt;&lt;a href=&quot;#fn:3&quot; class=&quot;footnote&quot; rel=&quot;footnote&quot;&gt;3&lt;/a&gt;&lt;/sup&gt;。它代表的是数据中隐藏的、可以通过有限算力提取出来的高级规律、长程依赖和底层逻辑。&lt;/p&gt;

&lt;p&gt;普通人面对庞杂的世界，往往被困在雪花屏前。他们贪婪地刷着短视频，阅读着碎片化的八卦，大脑全速运转去处理那些极高的时间受限熵。他们以为自己获得了海量信息，大脑却依然空空如也。&lt;/p&gt;

&lt;p&gt;高手面对同样的世界，采取的策略截然不同。他们极度克制地过滤掉随机噪音，把所有宝贵的算力用来从中提取 Epiplexity。Nassim Nicholas Taleb 在经典著作 &lt;em&gt;Fooled by Randomness&lt;/em&gt; 中，曾用概率论揭示过类似的道理。Taleb 曾提到，他几乎不看社交媒体，甚至不读报纸。&lt;/p&gt;

&lt;p&gt;如果你每天观察自己的投资组合，你会看到无数微小的波动。在 Taleb 看来，这些波动中 99% 都是噪音，只有 1% 是信号。那些频繁查看报价的投资者，实际上是把算力浪费在了捕捉随机漫步的噪声上。Taleb 认为，信息的质量随着观测频率的增加而急剧下降。这种高频的、杂乱的信息，正是论文中所说的时间受限熵。&lt;/p&gt;

&lt;p&gt;只有当你把观测的时间跨度拉长，从每天看一次改为每十年看一次，那些细碎的噪音才会自动湮灭。剩下的、真正改变命运的结构性趋势，才是 Epiplexity。高手深谙此道。他们主动放弃对即时反馈的关注，通过与外界保持距离，强行降低大脑处理高熵信息的负荷。这种克制，是提取底层结构的必要代价。&lt;/p&gt;

&lt;p&gt;论文中有一个非常绝妙的实验结论。研究人员对比了语言文本数据和高清图像数据，发现文本数据包含了极高的 Epiplexity。尽管一张高清图片（例如 CIFAR-5M 数据集中的图片）在计算机里占据的字节数极大，但其中超过 99% 的信息都是毫无规律的像素级随机噪声。文本的内容结构则完全不同，每一个词的排列都蕴含着极强的逻辑关联和抽象概念。一段经典名著的字节数可能远不如一张低像素的风景照，文本含有的 Epiplexity 却是图像的成百上千倍。&lt;/p&gt;

&lt;p&gt;这也完美解释了为什么当前的人工智能革命是由预训练的大语言模型（LLM）引领的。在海量文本上预训练的模型，能够涌现出惊人的跨领域泛化能力，它们甚至能零样本去解决复杂的逻辑推理题。仅仅看图的模型却极难做到这一点。&lt;/p&gt;

&lt;p&gt;生活也是如此。高密度的深度阅读、系统性的硬核思考，本质上就是在高效提取高 Epiplexity 的内容。单纯追逐短平快的瞬时感官刺激，充其量只是在大口吞咽庞大而无用的随机熵。&lt;/p&gt;

&lt;h2 id=&quot;二计算能力的诅咒与涌现的奇迹&quot;&gt;二、计算能力的诅咒与「涌现」的奇迹&lt;/h2&gt;

&lt;p&gt;这篇论文最精彩的洞见，在于彻底揭示了「算力受限」的核心价值。&lt;/p&gt;

&lt;p&gt;我们常常抱怨自己记忆力不够好、大脑处理信息的速度不够快。我们总幻想如果大脑能像超级计算机一样过目不忘、瞬间计算，我们就能无往不利。这篇论文用严格的数学证明证明了：唯有算力受限，才能逼迫一个系统产生「涌现」（Emergence）现象。&lt;/p&gt;

&lt;blockquote&gt;
  &lt;p&gt;&lt;strong&gt;真正的智慧，恰恰诞生于局限之中。&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;设想一下著名的 Conway 生命游戏（Conway’s Game of Life）。在这个庞大的二维网格里，黑白方块仅仅根据极其简单的几条相邻规则进行繁衍或死亡。如果让前文提到的拉普拉斯妖去预测未来，它毫无压力。它只需要严格按照基础规则，一步一步去推演每一个网格的微观状态。在超级计算体系的眼中，整个世界只有枯燥的 0 和 1。&lt;/p&gt;

&lt;p&gt;人类的大脑算力存在绝对瓶颈。我们无法在一秒钟内计算几万个网格的复杂演化。为了能够预测生命游戏未来的走向，人类被迫发明了一套全新的高级抽象词汇。我们观察到了特定形状的方块组合，并将它们命名为「滑翔机」（gliders）、「稳定块」（static blocks）和「振荡器」（oscillators）。我们甚至进一步总结出了滑翔机移动的固定速度，以及它们彼此碰撞时的宏观规律。&lt;/p&gt;

&lt;p&gt;在这个奇妙的过程中，无限算力的机器只看到了局部规则，算力受限的人类却敏锐地看到了「结构」。这种为了克服自身算力不足而提炼出的高级概念组合，正是我们需要获取的 Epiplexity。&lt;/p&gt;

&lt;p&gt;你可以把这种现象映射到细胞自动机（Elementary Cellular Automata）的实验中。研究人员让大语言模型去学习各种自动机的演化规则。像 Rule 15 这种规则过于简单，生成的画面全是周期性重复的图案，毫无学习价值。像 Rule 30 这种规则生成的画面则极其混乱，充满了伪随机信息，模型算力耗尽也毫无建树。只有像 Rule 54 这种处于混沌边缘地带的复杂规则，既包含了变化，又隐藏着宏观的几何逻辑，模型在学习它时提取到了极高的 Epiplexity。&lt;/p&gt;

&lt;p&gt;大语言模型的这一举动，极其生动地展示了什么是真实的学习过程。&lt;/p&gt;

&lt;p&gt;如果你试图死记硬背工作中的所有细枝末节，记住每一个客户说过的一字一句，记住每一行代码的微小变动，你仅仅是在把自己降级为一块低效的机械硬盘。而真正的高手会坦然接受自身大脑容量和处理速度的物理局限。他们主动放弃对微观变量的机械穷举，转而死磕事物背后的宏观模式与底层规律。&lt;/p&gt;

&lt;p&gt;你记不住所有的树叶，所以你发明了「树」这个概念。你算不出所有的价格波动，所以你总结出了「周期」和「均值回归」。&lt;mark&gt;物理局限不仅没有锁死人类的智慧，它恰恰是人类迈向顶层认知的核心驱动力。&lt;/mark&gt;&lt;/p&gt;

&lt;h2 id=&quot;三因式分解的痛苦是重塑大脑的捷径&quot;&gt;三、因式分解的痛苦，是重塑大脑的捷径&lt;/h2&gt;

&lt;p&gt;在经典的信息论中，还有一个长期被奉为圭臬的假设：信息的总量与观测时的分解顺序毫无关联。先观察要素 A 再观察要素 B，和先观察 B 再观察 A，你最终获得的总信息量理应完全一致。这在数学上被称为信息的对称性。&lt;/p&gt;

&lt;p&gt;真实世界的反馈完全粉碎了这个美好的错觉。论文作者通过一项极其硬核的国际象棋实验，彻底打碎了这层理论滤镜，揭示了认知升级的第三个秘密。&lt;/p&gt;

&lt;p&gt;研究人员将成千上万局国际象棋（Lichess 数据集）的专业棋谱喂给 AI 模型进行严苛的训练。他们特意使用了两种截然不同的数据排列方式：&lt;/p&gt;

&lt;ol&gt;
  &lt;li&gt;&lt;strong&gt;顺向逻辑。&lt;/strong&gt; 模型先看到一长串的走棋序列（比如白方走 E4，黑方走 E5），最后再让模型看一眼最终的棋盘状态布局。&lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;逆向逻辑。&lt;/strong&gt; 模型先看到最终的定格棋盘状态，然后再让它去预测和反推之前复杂的走棋序列。&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;实验结果非常令人震撼。对模型来说，第二种逆向训练法的预测难度飙升。根据经典理论，两组数据包含的绝对信息量是一模一样的，只是顺序变了。但在实际训练中，逆向预测强行迫使模型提取到了更为丰厚的 Epiplexity。&lt;/p&gt;

&lt;p&gt;在随后的未知任务测试中，这种差异展现得淋漓尽致。研究人员让这两个模型去完成一项它们从未见过的任务：给一个陌生的棋局打分，评估双方的优劣势（所谓的 Centipawn 评估）。经历了逆向训练的模型展现出了极强的分布外（OOD）泛化能力，得分远超第一种顺向模型。&lt;/p&gt;

&lt;p&gt;为什么会这样？答案藏在计算的非对称性里。因为艰难的逆向推导逼迫模型彻底放弃了表面的统计学死记硬背，它必须在大脑的深层神经网络中，建立起对棋盘整体局势极其深刻的内部表征。&lt;/p&gt;

&lt;p&gt;由因推果往往是一条平坦的高速公路。给你一步步的走法，顺水推舟推导出最终盘面，这只需要简单的规则叠加。这就好比你在看一本侦探小说，作者按照时间顺序平铺直叙，你毫不费力就能读到结尾。&lt;/p&gt;

&lt;p&gt;由果推因则是一条崎岖的蜀道。看到最终棋盘，你要在脑海里穷举无数种可能的历史路径，还要反向推演每一步的战术意图。由于算力受限，模型不可能使用暴力穷举法去倒推。艰难的逆向推导逼迫模型彻底放弃了表面的统计学死记硬背，它必须在大脑的深层神经网络中，建立起对棋盘整体局势、棋子价值和高阶战术极其深刻的内部表征。这套内部表征，就是极其珍贵的 Epiplexity。&lt;/p&gt;

&lt;p&gt;我们在现实生活中学习任何一项新技能时，同样面临这两种截然不同的路径选择。&lt;/p&gt;

&lt;p&gt;顺向学习就像是坐在教室里听老师讲课，顺着前人铺好的现成推导过程一路滑行下来，听着高管分享成功经验，一切都显得无比丝滑。你感觉自己什么都听懂了，大脑深处却没有建立起任何深刻的认知回路。这种知识是脆弱的，一旦遇到陌生的跨领域问题，顺向积累的经验瞬间崩塌。&lt;/p&gt;

&lt;p&gt;真正的刻意练习，必定包含这种逆向的、极其困难的「因式分解」。丢给你一个已经大获成功的商业案例，没有任何背景提示，让你孤立无援地反推创始人当初面临的生死抉择。拿给你一个业界顶尖的工业产品，让你从零开始逆向工程它的核心设计思路。&lt;/p&gt;

&lt;p&gt;这条路布满荆棘，它会急剧消耗你的心智与认知资源，让你感到极度的挫败。正因为如此，它能最大程度地打破你现有的神经连接，把冰冷的信息转化为你大脑中鲜活的 Epiplexity。高手都在日复一日地刻意给自己制造这种心智训练上的「剧烈不适感」。&lt;/p&gt;

&lt;blockquote&gt;
  &lt;p&gt;&lt;strong&gt;舒适区里永远只有低质的熵，只有在极其费力的逆向解构中，你才能提炼出智慧的黄金。&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2 id=&quot;四闭关锁国与算法的左脚踩右脚&quot;&gt;四、闭关锁国与算法的左脚踩右脚&lt;/h2&gt;

&lt;p&gt;我们经常听到一种笃定的说法：一个人如果不去积极接触外界的新信息，不保持对前沿资讯的极度饥渴，他就永远不可能产生新的认知跃迁。经典信息论中的「数据处理不等式」（Data Processing Inequality）也在用冷冰冰的数学语言表达同样的观点：对现有数据进行完全确定性的计算变换，绝对不可能凭空增加任何一丁点信息量。&lt;/p&gt;

&lt;p&gt;但是 Google 的 DeepMind 开发的 AlphaZero，或者说 0-shot learning，是对这一铁律的一个有力的反击。&lt;/p&gt;

&lt;p&gt;AlphaZero 在训练之初，仅仅被赋予了最基础的国际象棋规则。它完全拒绝输入任何人类高手留下的海量棋谱，没有任何外部知识的输入。它仅仅通过在封闭系统内，左右互搏，不知疲倦地进行自我对弈。几天之后，它演化出了极其深奥、甚至让整个人类国际象棋界都叹为观止的全新战略。它甚至发明了人类棋手几百年都没想到的弃子开局法。&lt;/p&gt;

&lt;p&gt;根据数据处理不等式，没有任何外部新数据的注入，AlphaZero 的系统总信息量应该始终为零。那个庞大的、包含了几千万个参数的神经网络里，装着的极其复杂的战略思维，究竟是从哪里凭空冒出来的？&lt;/p&gt;

&lt;p&gt;论文研究者给出了一个极度精辟的理论解释。当我们惊叹于 AlphaZero 学到了「新知识」时，我们谈论的完全脱离了 Shannon 意义上的信息量范畴，其本质依然是 Epiplexity。&lt;/p&gt;

&lt;p&gt;数据处理不等式有一个隐藏的致命前提：它假定观察者拥有无限算力。对于无限算力的上帝来说，国际象棋的规则和国际象棋的最优解，在信息量上是完全等价的。有了规则，最优解就已经注定存在于那里了。&lt;/p&gt;

&lt;p&gt;真实世界的观察者算力是极其有限的。规则虽然简单，但从规则推导出高阶战略，需要跨越一条极度宽广的计算鸿沟。对于一个计算资源受限的系统来说，通过投入大量的「计算」过程，完全可以把极其确定性的、看似毫无增量信息的基本公理，硬生生地转化为极具实战价值的结构化信息。&lt;/p&gt;

&lt;blockquote&gt;
  &lt;p&gt;&lt;strong&gt;计算的动作本身，就是在源源不断地创造新知识。&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;这也完美解释了为什么合成数据（Synthetic Data）能够让现代大模型变得更聪明，而不是所谓的 “Garbage in, garbage out”。模型用自己生成的看似没有新意的数据去继续训练自己，这种「左脚踩右脚」上天的做法，在古典理论看来是荒谬的。在算力受限的理论框架下，模型通过推演和生成，实际上是在将潜藏的深层结构显性化。&lt;/p&gt;

&lt;p&gt;将这个原理投射到人类历史上，你会发现那些最顶级的思想家，往往也有这样的经历。以 Isaac Newton 为例。当年为了躲避伦敦的严重瘟疫，牛顿把自己关在偏僻的伍尔索普庄园。长达一年多的时间里，他彻底切断了与外界学术圈的交流。没有任何新信息的输入，仅凭极其有限的几个基础物理学直觉和极其简单的数学公理，牛顿在自己的头脑中疯狂推演，最终构建出了极其庞大且严密的经典力学体系和微积分基本定理。&lt;/p&gt;

&lt;p&gt;王阳明也如此。《明史》载：&lt;/p&gt;

&lt;blockquote&gt;
  &lt;p&gt;（阳明）谪龙场，穷荒无书，日绎旧闻。忽悟格物致知，当自求诸心，不当求诸事物，喟然曰：「道在是矣。」遂笃信不疑。其为教，专以致良知为主。谓宋周、程二子后，惟象山陆氏简易直捷，有以接孟氏之传。而朱子「集注」、「或问」之类，乃中年未定之说。学者翕然从之，世遂有「阳明学」云。&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;现代社会的绝大多数人过分迷信「获取新信息」的神奇功效。他们患有严重的信息错失恐惧症（Fearing of Missing Out, FOMO），仿佛每天强迫自己刷完几百篇干货满满的行业报告、听完几十个观点前卫的播客，就能自动进化得更加聪明。&lt;/p&gt;

&lt;p&gt;这是一种本末倒置的错觉。真正的知识壁垒，极度依赖于深度的内在计算。当你费尽心思掌握了足够优秀的第一性原理之后，你最需要做的动作是果断关上房门，彻底切断外部世界汹涌的随机熵流，用你自己的大脑去进行残酷的自我博弈。&lt;/p&gt;

&lt;p&gt;去疯狂推演、去反复计算、去激烈碰撞。去思考一个极简法则在不同极端场景下的变体。这种看似枯燥的确定性内部重组，绝对可以为你淬炼出极其惊人的结构化智慧。&lt;/p&gt;

&lt;p&gt;&lt;mark&gt;知识的质变，永远发生在深刻的内部计算中，绝不发生在浮躁的外部检索里。&lt;/mark&gt;&lt;/p&gt;

&lt;h2 id=&quot;结语&quot;&gt;结语&lt;/h2&gt;

&lt;p&gt;回到我们文章最初的那个思考。真实世界的信息学，本质上是一门关于认知资源极限分配的硬核科学。&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;因为我们的物理生命和脑力算力是极其有限的，&lt;/strong&gt;我们要坚决拒绝把宝贵的算力浪费在那些看似热闹、实则高熵的随机事件上。新闻的头条、短期的股价波动、八卦绯闻，这些东西只会耗尽你的时间受限熵。你的目标是寻找那些历经时间考验、具备深层逻辑和长程依赖的知识，去疯狂提取 Epiplexity。&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;因为我们的物理生命和脑力算力是极其有限的，&lt;/strong&gt;我们要彻底放弃做一个试图记住所有细节的人肉照相机，勇敢地拥抱你的「算力不足」。正因为你记不住，你才会被迫去寻找事物背后的宏观模式，去总结规律，去创造抽象概念。局限性，是你获得涌现智慧的最强催化剂。&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;因为我们的物理生命和脑力算力是极其有限的，&lt;/strong&gt;我们要刻意且勇敢地选择那条逆向的、极为难走的推理小径，去逼迫大脑完成真正的底层升级。警惕那些嚼碎了喂给你的顺向知识。去拆解、去逆向工程、去由果推因。在极度烧脑的不适感中，你的大脑神经元正在发生实质性的重连。&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;因为我们的物理生命和脑力算力是极其有限的，&lt;/strong&gt;我们要停止无休止的外部信息摄入。留出大段的空白时间，用你已经掌握的第一性原理，在头脑中进行深刻的演算和自我博弈。通过深度的内部计算，你完全可以凭空创造出属于你自己的新知识体系。&lt;/p&gt;

&lt;p&gt;所谓高手，头脑无比清醒地知道自己只不过是一个「算力受限」的凡躯。他们绝不奢求去充当那个全知全能的拉普拉斯妖。他们只是在这个充斥着无穷无尽噪音的荒芜宇宙里，心无旁骛、坚定不移地执行着提取 Epiplexity 的终极生存算法。&lt;/p&gt;

&lt;hr /&gt;

&lt;div class=&quot;footnotes&quot; role=&quot;doc-endnotes&quot;&gt;
  &lt;ol&gt;
    &lt;li id=&quot;fn:1&quot; role=&quot;doc-endnote&quot;&gt;
      &lt;p&gt;对于非物理学背景的读者，「拉普拉斯妖」（Laplace’s Demon）是科学史上一个著名的思维实验，由法国数学家皮埃尔-西蒙·拉普拉斯（Pierre-Simon Laplace）于 1814 年提出。拉普拉斯假设，如果宇宙中存在一个智能生物，能够掌握某一瞬时所有物质的精确位置与动量，并具备处理这些数据的超凡算力，那么根据经典力学的因果律，整个宇宙的过去与未来对其而言都将是确定的。这一概念是「机械决定论」的极致体现，暗示宇宙如同精密运转的钟表，一切演化皆可预见。然而，20 世纪量子力学的发展打破了这一幻想。海森堡提出的不确定性原理（Uncertainty Principle）从底层逻辑上证明，微观粒子的位置与动量无法同时被精确测量。这意味着即便存在所谓的「妖」，它也无法获得推演未来的初始数据，从而在科学层面宣告了这一全知模型的失效。 &lt;a href=&quot;#fnref:1&quot; class=&quot;reversefootnote&quot; role=&quot;doc-backlink&quot;&gt;&amp;#8617;&lt;/a&gt;&lt;/p&gt;
    &lt;/li&gt;
    &lt;li id=&quot;fn:2&quot; role=&quot;doc-endnote&quot;&gt;
      &lt;p&gt;没错，Anthropic 公司的 AI 产品就是用的 Shannon 的名字。个人认为 A 司是一个有 taste 的公司。 &lt;a href=&quot;#fnref:2&quot; class=&quot;reversefootnote&quot; role=&quot;doc-backlink&quot;&gt;&amp;#8617;&lt;/a&gt;&lt;/p&gt;
    &lt;/li&gt;
    &lt;li id=&quot;fn:3&quot; role=&quot;doc-endnote&quot;&gt;
      &lt;p&gt;对于熟悉词源学的读者。Epiplexity, literally，指的是「在原有复杂性之上的复杂性」。&lt;/p&gt;

      &lt;ul&gt;
        &lt;li&gt;&lt;strong&gt;前缀 Epi-（ἐπί）：&lt;/strong&gt;在希腊语中意为「在……之上」、「附加」或「外层」。通常表示一个更高的维度、叠加的层次或是在基础结构之上的衍生。&lt;/li&gt;
        &lt;li&gt;&lt;strong&gt;词根 -plexity：&lt;/strong&gt;源自拉丁语 &lt;em&gt;plectere&lt;/em&gt;（编织、缠绕）。它与 Complex（复杂）同源，指的是多个元素交织在一起，形成一个难以拆分的整体。&lt;/li&gt;
      &lt;/ul&gt;
      &lt;p&gt;&lt;a href=&quot;#fnref:3&quot; class=&quot;reversefootnote&quot; role=&quot;doc-backlink&quot;&gt;&amp;#8617;&lt;/a&gt;&lt;/p&gt;
    &lt;/li&gt;
  &lt;/ol&gt;
&lt;/div&gt;
</content>
  </entry>
  
  <entry>
    <title>Does Less Compute Mean Higher Cognition?</title>
    <link href="https://mochiaochen.github.io/en/writing/2026/07/lower-compute-higher-cognition/" rel="alternate" type="text/html"/>
    <published>2026-07-22T15:30:00+08:00</published>
    <updated>2026-07-22T15:30:00+08:00</updated>
    <id>https://mochiaochen.github.io/en/writing/2026/07/lower-compute-higher-cognition-en</id>
    <content type="html" xml:base="https://mochiaochen.github.io/en/writing/2026/07/lower-compute-higher-cognition/">&lt;p&gt;Suppose there were a spectre in the universe called “Laplace’s Demon.”&lt;sup id=&quot;fnref:1-en&quot; role=&quot;doc-noteref&quot;&gt;&lt;a href=&quot;#fn:1-en&quot; class=&quot;footnote&quot; rel=&quot;footnote&quot;&gt;1&lt;/a&gt;&lt;/sup&gt; It possesses infinite computational power and knows the momentum and position of every atom in the universe. You might think that such a being must possess the highest wisdom in the universe, able to see through.&lt;/p&gt;

&lt;p&gt;The reality is completely contrary to our expectations. Mathematics and computer science tell us a brutal truth: &lt;mark&gt;infinite compute means that it has no need for what humans define as “intelligence.”&lt;/mark&gt; In the eyes of Laplace’s Demon, the world contains no “apples,” no “countries,” and no “love”—only a collection of atoms moving tediously according to physical laws. It does not need to summarise patterns, because it can brute-force everything directly.&lt;/p&gt;

&lt;p&gt;A striking paper titled &lt;em&gt;From Entropy to Epiplexity&lt;/em&gt; recently appeared on arXiv. Several researchers from Carnegie Mellon University and New York University re-examined classical information theory and proposed a deeply counterintuitive conclusion: &lt;strong&gt;if you want to learn a genuinely useful system of knowledge from the world and build what we call higher cognition, you must be a computationally bounded observer.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Although the paper uses rigorous mathematical formulas and cryptographic concepts throughout to discuss large language models (LLMs) and machine learning, it actually reveals a profound algorithm for life.&lt;/p&gt;

&lt;h2 id=&quot;i-shannons-blind-spot-and-the-curse-of-television-static&quot;&gt;I. Shannon’s blind spot and the curse of television static&lt;/h2&gt;

&lt;p&gt;In classical information theory, information is measured in bits. Claude Shannon&lt;sup id=&quot;fnref:2-en&quot; role=&quot;doc-noteref&quot;&gt;&lt;a href=&quot;#fn:2-en&quot; class=&quot;footnote&quot; rel=&quot;footnote&quot;&gt;2&lt;/a&gt;&lt;/sup&gt; tells us that the essence of information is “the elimination of uncertainty.” The less predictable and more chaotic an event is, the more information it contains. This is what we call information entropy. By this standard, we are mercilessly bombarded with enormous volumes of information every day.&lt;/p&gt;

&lt;p&gt;Some may say that we need more information to find the “right” path through life. But imagine an old television with no antenna plugged in, its screen filled with dense static. The image is completely random: you can never predict whether the next pixel will be black or white. According to Shannon’s theory, and the later complexity theory of the Soviet mathematician Kolmogorov, that field of static contains an enormous amount of information—nearly the theoretical maximum.&lt;/p&gt;

&lt;blockquote&gt;
  &lt;p&gt;Yet no normal person would stare at television static all day. There is no “meaning” in it whatsoever.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;To fill this gap, the paper’s authors introduce two new core concepts: &lt;strong&gt;Time-bounded Entropy&lt;/strong&gt; and &lt;strong&gt;Epiplexity&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Time-bounded Entropy refers to unpredictable content that, given limited time and compute, appears entirely random and impossible to make sense of. Most of life’s trivia, emotional venting on social media, and the stock market’s high-frequency daily fluctuations all constitute extremely high Time-bounded Entropy for an ordinary person. No matter how much effort you put into tracking them, you cannot extract any reusable pattern. You are simply spending your life calculating something akin to a cryptographic pseudorandom sequence.&lt;/p&gt;

&lt;p&gt;By contrast, Epiplexity can be translated precisely as “structural complexity.”&lt;sup id=&quot;fnref:3-en&quot; role=&quot;doc-noteref&quot;&gt;&lt;a href=&quot;#fn:3-en&quot; class=&quot;footnote&quot; rel=&quot;footnote&quot;&gt;3&lt;/a&gt;&lt;/sup&gt; It represents the higher-order patterns, long-range dependencies, and underlying logic hidden in data that can be extracted with finite compute.&lt;/p&gt;

&lt;p&gt;Confronted with the world’s complexity, ordinary people often become trapped in front of the static. They greedily scroll through short videos and consume fragmented gossip, their brains running at full speed to process extremely high Time-bounded Entropy. They believe they have acquired a vast amount of information, yet their minds remain empty.&lt;/p&gt;

&lt;p&gt;Experts adopt an entirely different strategy toward the same world. With extreme restraint, they filter out random noise and devote all their precious compute to extracting Epiplexity. In his classic &lt;em&gt;Fooled by Randomness&lt;/em&gt;, Nassim Nicholas Taleb used probability theory to reveal a similar principle. Taleb has said that he barely looks at social media and does not even read newspapers.&lt;/p&gt;

&lt;p&gt;If you observe your portfolio every day, you will see countless tiny fluctuations. In Taleb’s view, 99% of those movements are noise and only 1% is signal. Investors who constantly check quotations are wasting compute trying to capture the noise of a random walk. Taleb argues that the quality of information falls sharply as the frequency of observation rises. This high-frequency, disorderly information is exactly what the paper calls Time-bounded Entropy.&lt;/p&gt;

&lt;p&gt;Only when you lengthen the observation window—from once a day to once a decade—does the fine-grained noise automatically vanish. What remains, the structural trends that truly change one’s destiny, is Epiplexity. Experts understand this deeply. They actively give up their fixation on instant feedback and keep their distance from the outside world, forcibly reducing the burden that high-entropy information places on the brain. Such restraint is the necessary cost of extracting underlying structure.&lt;/p&gt;

&lt;p&gt;The paper contains a brilliant experimental finding. Researchers compared language-text data with high-definition image data and found that text contains extremely high Epiplexity. Although a high-definition image—one from the CIFAR-5M dataset, for example—occupies a huge number of bytes in a computer, more than 99% of its information is irregular pixel-level random noise. Text has a completely different structure: the arrangement of every word embodies strong logical relationships and abstract concepts. A passage from a classic may occupy far fewer bytes than a low-resolution landscape photograph, yet contain hundreds or thousands of times as much Epiplexity.&lt;/p&gt;

&lt;p&gt;This also perfectly explains why today’s AI revolution has been led by pretrained large language models. Models pretrained on massive bodies of text can develop astonishing cross-domain generalisation, even solving complex logical reasoning problems zero-shot. Models that only look at images find this extraordinarily difficult.&lt;/p&gt;

&lt;p&gt;Life works the same way. Dense, deep reading and systematic, rigorous thought are essentially efficient ways to extract content rich in Epiplexity. Chasing nothing but brief, fast sensory stimulation amounts to gulping down enormous quantities of useless random entropy.&lt;/p&gt;

&lt;h2 id=&quot;ii-the-curse-of-computation-and-the-miracle-of-emergence&quot;&gt;II. The curse of computation and the miracle of “emergence”&lt;/h2&gt;

&lt;p&gt;The paper’s most compelling insight is its complete exposition of the central value of being “compute-bounded.”&lt;/p&gt;

&lt;p&gt;We often complain that our memory is not good enough and that our brains process information too slowly. We imagine that if our brains could remember everything and calculate instantly like supercomputers, nothing could stand in our way. With rigorous mathematical proof, the paper proves that only computational constraints can force a system to produce the phenomenon of “emergence.”&lt;/p&gt;

&lt;blockquote&gt;
  &lt;p&gt;&lt;strong&gt;True intelligence is born precisely from limitation.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Consider Conway’s Game of Life. Across its vast two-dimensional grid, black and white cells live or die according to only a few extremely simple rules about their neighbours. Laplace’s Demon would have no difficulty predicting its future. It would merely follow the basic rules and calculate, step by step, the microscopic state of every cell. In the eyes of a supercomputing system, the whole world contains nothing but tedious zeroes and ones.&lt;/p&gt;

&lt;p&gt;The human brain faces an absolute computational bottleneck. We cannot calculate the complex evolution of tens of thousands of cells in a second. To predict where the Game of Life is heading, people were forced to invent an entirely new vocabulary of higher abstractions. We observed combinations of cells with particular shapes and named them “gliders,” “static blocks,” and “oscillators.” We went on to summarise the fixed speed at which gliders move and the macroscopic laws governing their collisions.&lt;/p&gt;

&lt;p&gt;In this remarkable process, the machine with infinite compute sees only local rules, while computationally bounded humans perceive “structure.” These combinations of higher-order concepts, distilled to overcome our lack of compute, are precisely the Epiplexity we need to acquire.&lt;/p&gt;

&lt;p&gt;The same phenomenon can be mapped onto experiments with Elementary Cellular Automata. Researchers asked large language models to learn the evolutionary rules of different automata. A rule such as Rule 15 is too simple: its images consist entirely of periodic, repeating patterns with no learning value. Rule 30 produces images so chaotic and full of pseudorandom information that a model can exhaust its compute without achieving anything. Only a complex rule such as Rule 54, poised at the edge of chaos, both contains change and conceals a macroscopic geometric logic. In learning it, the model extracts high Epiplexity.&lt;/p&gt;

&lt;p&gt;This behaviour by large language models vividly demonstrates what genuine learning is.&lt;/p&gt;

&lt;p&gt;If you try to memorise every detail of your work, every word every client has spoken, and every minute change to every line of code, you are merely downgrading yourself into an inefficient mechanical hard drive. True experts calmly accept the physical limits of their brain capacity and processing speed. They actively abandon mechanical enumeration of microscopic variables and instead wrestle with the macroscopic patterns and underlying laws behind things.&lt;/p&gt;

&lt;p&gt;You cannot remember every leaf, so you invent the concept of a “tree.” You cannot calculate every price movement, so you summarise “cycles” and “mean reversion.” &lt;mark&gt;Physical limitation does not imprison human intelligence; it is the central force driving humanity toward higher cognition.&lt;/mark&gt;&lt;/p&gt;

&lt;h2 id=&quot;iii-the-pain-of-factorisation-is-a-shortcut-to-reshaping-the-brain&quot;&gt;III. The pain of factorisation is a shortcut to reshaping the brain&lt;/h2&gt;

&lt;p&gt;Classical information theory has another long-revered assumption: the total amount of information is unrelated to the order in which an observation is decomposed. Observe element A and then element B, or B and then A, and the total information you ultimately obtain should be identical. Mathematically, this is called the symmetry of information.&lt;/p&gt;

&lt;p&gt;Feedback from the real world completely shatters this beautiful illusion. Through a demanding chess experiment, the paper’s authors broke through this theoretical filter and revealed a third secret of cognitive advancement.&lt;/p&gt;

&lt;p&gt;The researchers rigorously trained AI models on professional records from tens of thousands of chess games in the Lichess dataset. They deliberately arranged the data in two radically different ways:&lt;/p&gt;

&lt;ol&gt;
  &lt;li&gt;&lt;strong&gt;Forward logic.&lt;/strong&gt; The model first saw a long sequence of moves—White plays E4, Black plays E5, for instance—and only at the end saw the final board position.&lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Reverse logic.&lt;/strong&gt; In a deeply counterintuitive arrangement, the model first saw the final frozen board position and then had to predict and infer the preceding complex sequence of moves.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The results were striking. For the model, the difficulty of prediction soared under the second, reverse-training method. Classical theory says that the two datasets contain exactly the same absolute amount of information; only the order has changed. In actual training, however, reverse prediction forcibly made the model extract richer Epiplexity.&lt;/p&gt;

&lt;p&gt;The difference became fully apparent in a subsequent test on an unfamiliar task. The researchers asked the two models to do something neither had ever seen: score an unfamiliar chess position and assess which side held the advantage, in a so-called centipawn evaluation. The reverse-trained model demonstrated extremely strong out-of-distribution (OOD) generalisation and scored far above the first, forward-trained model.&lt;/p&gt;

&lt;p&gt;Why? The answer lies in the asymmetry of computation. Difficult reverse inference forces a model to abandon superficial statistical memorisation entirely. Deep inside its neural network, it must construct a profound internal representation of the overall position on the board.&lt;/p&gt;

&lt;p&gt;Moving from cause to effect is often a flat highway. Given the moves one by one, you can flow naturally toward the final position through simple accumulation of rules. It is like reading a detective novel narrated chronologically: you effortlessly reach the ending.&lt;/p&gt;

&lt;p&gt;Moving from effect to cause is a rugged mountain road. Seeing the final board, you must enumerate countless possible historical paths in your mind and work backwards through the tactical intention behind every move. Because its compute is limited, the model cannot brute-force the reconstruction. Difficult reverse inference compels it to abandon surface-level statistical memorisation and build, deep within its neural network, a profound internal representation of the overall position, piece values, and high-level tactics. This internal representation is precious Epiplexity.&lt;/p&gt;

&lt;p&gt;When learning any new skill in real life, we face the same choice between two radically different paths.&lt;/p&gt;

&lt;p&gt;Forward learning is like sitting in a classroom, gliding along a ready-made derivation laid down by others and listening to executives share the secrets of their success. Everything feels perfectly smooth. You feel that you understand it all, yet no deep cognitive circuit has formed in your brain. Such knowledge is fragile: the moment it encounters an unfamiliar cross-domain problem, experience accumulated in the forward direction collapses.&lt;/p&gt;

&lt;p&gt;True deliberate practice must contain this reverse, extraordinarily difficult “factorisation.” Take a wildly successful business case, stripped of any background hints, and infer in isolation the life-or-death choices its founder originally faced. Take a leading industrial product and reverse-engineer its core design logic from scratch.&lt;/p&gt;

&lt;p&gt;This road is covered in thorns. It rapidly consumes mental and cognitive resources and leaves you profoundly frustrated. Precisely for that reason, it can break your existing neural connections to the greatest possible degree and convert cold information into living Epiplexity inside your brain. Day after day, experts deliberately create this “intense discomfort” in their mental training.&lt;/p&gt;

&lt;blockquote&gt;
  &lt;p&gt;&lt;strong&gt;The comfort zone contains nothing but low-grade entropy. Only through immensely demanding reverse decomposition can you refine the gold of wisdom.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2 id=&quot;iv-closing-the-country-and-an-algorithm-pulling-itself-up-by-its-bootstraps&quot;&gt;IV. Closing the country and an algorithm pulling itself up by its bootstraps&lt;/h2&gt;

&lt;p&gt;We often hear a confident claim: if a person does not actively engage with new information from the outside world, and does not maintain an intense hunger for the latest developments, they can never make a new cognitive leap. The Data Processing Inequality of classical information theory expresses the same view in cold mathematical language: applying a fully deterministic computational transformation to existing data can never create even the slightest additional amount of information out of nothing.&lt;/p&gt;

&lt;p&gt;But AlphaZero, developed by Google’s DeepMind—or, one might say, 0-shot learning—is a powerful rebuttal to this iron law.&lt;/p&gt;

&lt;p&gt;At the beginning of training, AlphaZero was given only the most basic rules of chess. It rejected the massive archive of game records left by human masters and received no external knowledge. Inside a closed system, it simply played tirelessly against itself. A few days later, it had evolved extraordinarily deep new strategies that astonished the entire human chess world. It even invented sacrificial openings that human players had not conceived of in hundreds of years.&lt;/p&gt;

&lt;p&gt;According to the Data Processing Inequality, with no injection of new external data, the total information in AlphaZero’s system should have remained zero. Where, then, did the enormously complex strategic thought housed in that vast neural network, with its tens of millions of parameters, suddenly come from?&lt;/p&gt;

&lt;p&gt;The paper’s researchers offer an incisive theoretical explanation. When we marvel that AlphaZero has learned “new knowledge,” what we mean lies entirely outside information quantity in Shannon’s sense. Its essence remains Epiplexity.&lt;/p&gt;

&lt;p&gt;The Data Processing Inequality has a hidden, fatal premise: it assumes an observer with infinite compute. For a god of infinite computation, the rules of chess and the optimal solution to chess are equivalent in information content. Once the rules exist, the optimal solution is already destined to be there.&lt;/p&gt;

&lt;p&gt;Observers in the real world have extremely limited compute. The rules may be simple, but deriving high-level strategy from them requires crossing an immensely wide computational gulf. By investing an enormous amount of “computation,” a resource-bounded system can forcibly transform highly deterministic basic axioms, apparently carrying no incremental information, into structured information of enormous practical value.&lt;/p&gt;

&lt;blockquote&gt;
  &lt;p&gt;&lt;strong&gt;The act of computation itself continually creates new knowledge.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;This also perfectly explains why Synthetic Data can make modern large models smarter, rather than producing “garbage in, garbage out.” A model continues training on apparently unoriginal data that it generated itself. From the perspective of classical theory, this attempt to lift itself by its own bootstraps is absurd. Under a computationally bounded framework, however, the model uses inference and generation to make latent deep structure explicit.&lt;/p&gt;

&lt;p&gt;Project this principle onto human history and you will find that the greatest thinkers often had similar experiences. Consider Isaac Newton. To escape the severe plague in London, Newton shut himself away at the remote Woolsthorpe Manor. For more than a year, he completely severed contact with the outside academic world. With no new information coming in, relying only on a few basic physical intuitions and very simple mathematical axioms, Newton calculated furiously inside his own mind and ultimately constructed the vast, rigorous system of classical mechanics and the fundamental theorem of calculus.&lt;/p&gt;

&lt;p&gt;Wang Yangming was the same. The &lt;em&gt;History of Ming&lt;/em&gt; records:&lt;/p&gt;

&lt;blockquote&gt;
  &lt;p&gt;(Yangming) was exiled to Longchang, a desolate place without books, and daily worked through what he had learned before. Suddenly he realised that the investigation of things and the extension of knowledge should be sought within one’s own mind, not in external objects. He sighed, “The Way is here.” Thereafter he believed without doubt. His teaching centred on extending innate knowledge. He held that after Zhou and the two Cheng brothers of the Song, only Lu Xiangshan’s simple and direct approach continued the transmission from Mencius, while Zhu Xi’s &lt;em&gt;Collected Commentaries&lt;/em&gt;, &lt;em&gt;Questions and Answers&lt;/em&gt;, and similar works were unsettled views from Zhu’s middle years. Scholars followed him in great numbers, and the world thus came to speak of the “Yangming School.”&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Most people in modern society place excessive faith in the magical power of “acquiring new information.” They suffer from a severe Fear of Missing Out (FOMO), as if forcing themselves to scroll through hundreds of information-packed industry reports and listen to dozens of podcasts with avant-garde opinions every day would automatically make them smarter.&lt;/p&gt;

&lt;p&gt;This is an illusion that puts the cart before the horse. Genuine barriers of knowledge depend intensely on deep internal computation. Once you have worked hard to master sufficiently strong first principles, the action you most need is to close the door decisively, cut off the raging stream of random entropy from the outside world, and use your own brain to wage a ruthless game against itself.&lt;/p&gt;

&lt;p&gt;Infer furiously, calculate repeatedly, and collide violently. Think through how a minimalist rule changes under different extreme scenarios. This apparently tedious, deterministic internal reorganisation can absolutely temper astonishing structural intelligence within you.&lt;/p&gt;

&lt;p&gt;&lt;mark&gt;Qualitative change in knowledge always happens through profound internal computation, never through restless external search.&lt;/mark&gt;&lt;/p&gt;

&lt;h2 id=&quot;conclusion&quot;&gt;Conclusion&lt;/h2&gt;

&lt;p&gt;Return to the thought with which this essay began. Information science in the real world is, in essence, a rigorous science of allocating extremely limited cognitive resources.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Because our physical lives and mental compute are extremely limited,&lt;/strong&gt; we must firmly refuse to waste precious compute on random events that look lively but are in fact high-entropy. News headlines, short-term stock-price movements, and gossip merely exhaust your Time-bounded Entropy. Your goal is to find knowledge tested by time, with deep logic and long-range dependencies, and relentlessly extract its Epiplexity.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Because our physical lives and mental compute are extremely limited,&lt;/strong&gt; we must completely abandon the attempt to become a human camera that remembers every detail, and bravely embrace our “lack of compute.” Because you cannot remember everything, you are forced to seek the macroscopic patterns behind things, summarise laws, and create abstract concepts. Limitation is the strongest catalyst for emergent intelligence.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Because our physical lives and mental compute are extremely limited,&lt;/strong&gt; we must deliberately and courageously choose the reverse and extraordinarily difficult path of reasoning, forcing the brain to undergo a genuine foundational upgrade. Be wary of forward knowledge that has been chewed up and fed to you. Decompose, reverse-engineer, and infer causes from effects. Amid the profoundly taxing discomfort, the neurons in your brain are undergoing substantive rewiring.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Because our physical lives and mental compute are extremely limited,&lt;/strong&gt; we must stop the endless intake of external information. Leave large stretches of empty time. Use the first principles you already possess to perform deep calculations and self-play inside your mind. Through deep internal computation, you can create an entirely new system of knowledge that belongs to you.&lt;/p&gt;

&lt;p&gt;The expert knows with absolute clarity that they are only a mortal vessel with limited compute. They never aspire to become the omniscient and omnipotent Laplace’s Demon. In this barren universe filled with endless noise, they simply devote themselves, unwaveringly and without distraction, to the ultimate survival algorithm: extracting Epiplexity.&lt;/p&gt;

&lt;hr /&gt;

&lt;div class=&quot;footnotes&quot; role=&quot;doc-endnotes&quot;&gt;
  &lt;ol&gt;
    &lt;li id=&quot;fn:1-en&quot; role=&quot;doc-endnote&quot;&gt;
      &lt;p&gt;For readers without a background in physics, “Laplace’s Demon” is a famous thought experiment in the history of science, proposed by the French mathematician Pierre-Simon Laplace in 1814. Laplace imagined an intelligent being that knew the precise position and momentum of all matter at a given instant and possessed extraordinary power to process the data. Under the causal laws of classical mechanics, the universe’s entire past and future would then be determined for it. The concept is the ultimate expression of mechanical determinism: a clockwork universe whose every development can be foreseen. The development of quantum mechanics in the twentieth century shattered this fantasy. Heisenberg’s Uncertainty Principle established at a fundamental level that a microscopic particle’s position and momentum cannot both be measured precisely. Even such a “demon” could therefore never obtain the initial data needed to infer the future, invalidating the omniscient model scientifically. &lt;a href=&quot;#fnref:1-en&quot; class=&quot;reversefootnote&quot; role=&quot;doc-backlink&quot;&gt;&amp;#8617;&lt;/a&gt;&lt;/p&gt;
    &lt;/li&gt;
    &lt;li id=&quot;fn:2-en&quot; role=&quot;doc-endnote&quot;&gt;
      &lt;p&gt;Yes, Anthropic’s AI product is named after Shannon. Personally, I think Company A is a company with taste. &lt;a href=&quot;#fnref:2-en&quot; class=&quot;reversefootnote&quot; role=&quot;doc-backlink&quot;&gt;&amp;#8617;&lt;/a&gt;&lt;/p&gt;
    &lt;/li&gt;
    &lt;li id=&quot;fn:3-en&quot; role=&quot;doc-endnote&quot;&gt;
      &lt;p&gt;For readers familiar with etymology: &lt;em&gt;Epiplexity&lt;/em&gt;, literally, means “complexity on top of existing complexity.”&lt;/p&gt;

      &lt;ul&gt;
        &lt;li&gt;&lt;strong&gt;The prefix Epi- (ἐπί):&lt;/strong&gt; Greek for “upon,” “added,” or “outer.” It commonly indicates a higher dimension, a superimposed layer, or something derived on top of a foundational structure.&lt;/li&gt;
        &lt;li&gt;&lt;strong&gt;The root -plexity:&lt;/strong&gt; From the Latin &lt;em&gt;plectere&lt;/em&gt;, “to weave” or “to entwine.” It shares an origin with &lt;em&gt;complex&lt;/em&gt; and refers to multiple elements woven together into a whole that is difficult to separate.&lt;/li&gt;
      &lt;/ul&gt;
      &lt;p&gt;&lt;a href=&quot;#fnref:3-en&quot; class=&quot;reversefootnote&quot; role=&quot;doc-backlink&quot;&gt;&amp;#8617;&lt;/a&gt;&lt;/p&gt;
    &lt;/li&gt;
  &lt;/ol&gt;
&lt;/div&gt;
</content>
  </entry>
  
</feed>
