<?xml version="1.0" encoding="utf-8"?>
<rss version="2.0" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/">
    <channel>
        <title>智汇AI</title>
        <link>http://easyai.fyi/</link>
        <description>每天看懂AI新闻、知识与工具，给产品经理、运营和AI初学者的人工智能科普站。</description>
        <lastBuildDate>Mon, 28 Sep 2026 07:35:54 GMT</lastBuildDate>
        <docs>https://validator.w3.org/feed/docs/rss2.html</docs>
        <generator>https://github.com/jpmonette/feed</generator>
        <language>zh-CN</language>
        <copyright>All rights reserved 2026, 智汇AI</copyright>
        <item>
            <title><![CDATA[Codex 用量与 AI-native 工作分化 - 2026-07-29]]></title>
            <link>http://easyai.fyi/article/follow-builders-ai-summary-2026-07-29</link>
            <guid>http://easyai.fyi/article/follow-builders-ai-summary-2026-07-29</guid>
            <pubDate>Wed, 29 Jul 2026 00:00:00 GMT</pubDate>
            <description><![CDATA[本轮 follow-builders 新增 10 个 X builders、23 条 tweet、0 篇 blog 和 0 期 podcast，已按输出全量沉淀。今天主线是 Codex / ChatGPT Work 的真实用量问题浮出水面，同时 AI-native IC、agent 工具链和创作工作流继续分化。]]></description>
            <content:encoded><![CDATA[<div id="notion-article" class="mx-auto overflow-hidden "><main class="notion light-mode notion-page notion-block-3ac5eac3813c8101a364f9a7a2b3c1a2"><div class="notion-viewport"></div><div class="notion-collection-page-properties"></div><h3 class="notion-h notion-h2 notion-h-indent-0 notion-block-cdf4bbd8a87641bea712f6436ba8cb2a" data-id="cdf4bbd8a87641bea712f6436ba8cb2a"><span><div id="cdf4bbd8a87641bea712f6436ba8cb2a" class="notion-header-anchor"></div><a class="notion-hash-link" href="#cdf4bbd8a87641bea712f6436ba8cb2a" title="今日主线"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">今日主线</span></span></h3><div class="notion-text notion-block-ff12410ff77149d7acb806f59af40fbf">今天的 follow-builders 内容集中在一个现实问题：AI 工具开始进入高强度使用后，能力、成本和工作方式会一起变化。Thibault Sottiaux 对 GPT-5.6 Sol 的 Codex / ChatGPT Work 用量做了长解释，承认更强的模型会做更多 tool calls、跑更复杂 workflow，也会让部分重度用户更快耗尽额度。Swyx 则把人才市场的分化说得更直白：会管理 agent 的 AI-native IC 和 player-coach 在牛市，而传统“head of X”式管理角色在承压。</div><h3 class="notion-h notion-h2 notion-h-indent-0 notion-block-b80d1a56c8dd48e1abd811c993221686" data-id="b80d1a56c8dd48e1abd811c993221686"><span><div id="b80d1a56c8dd48e1abd811c993221686" class="notion-header-anchor"></div><a class="notion-hash-link" href="#b80d1a56c8dd48e1abd811c993221686" title="重点解读"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">重点解读</span></span></h3><div class="notion-text notion-block-f7f401dbe06e4ba2bfe9448ca29a989b">第一条线是 Codex 用量从产品体验变成信任问题。Thibault Sottiaux 说，OpenAI 没有降低任何订阅计划的用量，但 Sol 比旧模型更愿意工作更久、调用更多工具、协调更多 subagents，导致一些任务消耗超出预期。OpenAI 预计典型 Sol 使用时长会提升约 18%，并会恢复此前临时暂停的五小时限制。这不是简单的“额度重置”，而是一个新阶段：agent 能力越强，用户越会用它处理复杂任务，计量、效率和透明解释就会变成产品体验的一部分。</div><div class="notion-text notion-block-5765e747954e40319df0a539073f939c">第二条线是 AI-native 工作能力的价值上升。Swyx 说，现在是 AI-native IC / player-coach 的大牛市，也是“heads of X”管理者的大熊市；他甚至用一句话概括：有一年管理 10 个 agent 的经验，可能比有十年管理 10 到 100 人的经验更有价值。这个判断很刺耳，但和 Dan Shipper 用 ChatGPT for Work voice mode 完成 Codex 史稿的案例放在一起看，方向是一致的：核心竞争力正在从“会不会用工具”升级到“能不能把 agent 放进真实工作流”。</div><div class="notion-text notion-block-0d17466e8404457da60de3543a5e6cdc">第三条线是安全和开源工具链继续往工程现场走。Thibault Sottiaux 还提到发布了用于发现、验证和修复代码安全漏洞的 CLI 和 TypeScript SDK，可以扫描仓库、检查变更、跟踪 findings，并放进 CI。Peter Steinberger 的一句“Serving large models is hard”虽然短，但也指向同一个现实：大家不再只讨论模型榜单，开始讨论部署、服务、成本、安全和工程化问题。</div><div class="notion-text notion-block-bf868d016f8543ee8a5a16b736115a02">第四条线是今天有不少非 AI 主线内容。Garry Tan 的几条主要围绕个人态度和旧金山政策，Amjad Masad 有 rare books 和数学证明搜索的联想，Aaron Levie 转发 Zuckerberg 的 AI vision 和 K3 weights 相关内容，但 JSON 里上下文有限。按全量沉淀要求，这些内容全部保留，但正文不额外补外部信息。</div><h3 class="notion-h notion-h2 notion-h-indent-0 notion-block-7204f9b4c4a5407c8c6e1da8966e411d" data-id="7204f9b4c4a5407c8c6e1da8966e411d"><span><div id="7204f9b4c4a5407c8c6e1da8966e411d" class="notion-header-anchor"></div><a class="notion-hash-link" href="#7204f9b4c4a5407c8c6e1da8966e411d" title="X Builders 全量记录"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">X Builders 全量记录</span></span></h3><h4 class="notion-h notion-h3 notion-h-indent-1 notion-block-6940924bddb24fb49b04cf0de197fa54" data-id="6940924bddb24fb49b04cf0de197fa54"><span><div id="6940924bddb24fb49b04cf0de197fa54" class="notion-header-anchor"></div><a class="notion-hash-link" href="#6940924bddb24fb49b04cf0de197fa54" title="1. Swyx @swyx"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">1. Swyx @swyx</span></span></h4><div class="notion-text notion-block-0b9fec2a99704370a8bdbedb5eeadd1c">Swyx 今天最有信息量的是对 AI-native 人才市场的判断：会亲手操盘 agent 的 IC / player-coach 正在走强，而传统“heads of X”管理角色正在承压。他还补充说自己只是观察趋势，不是说没有反例，并提到可以去听 Latent Space 里 Edison 的相关内容。另一条只有链接和引用关系，JSON 里没有足够上下文，这里不额外补充。</div><figure class="notion-asset-wrapper notion-asset-wrapper-embed notion-block-3c9c958e68394c87a5986263f7e73228"><div style="position:relative;display:flex;justify-content:center;align-self:center;width:100%;max-width:100%;flex-direction:column;height:320px"><iframe class="notion-asset-object-fit" src="https://platform.twitter.com/embed/Tweet.html?id=2082287480687272053" title="iframe embed" frameBorder="0" allowfullscreen="" loading="lazy" scrolling="auto"></iframe></div></figure><figure class="notion-asset-wrapper notion-asset-wrapper-embed notion-block-430609a42cca4ddfb7c22d201de95362"><div style="position:relative;display:flex;justify-content:center;align-self:center;width:100%;max-width:100%;flex-direction:column;height:320px"><iframe class="notion-asset-object-fit" src="https://platform.twitter.com/embed/Tweet.html?id=2082255848492183583" title="iframe embed" frameBorder="0" allowfullscreen="" loading="lazy" scrolling="auto"></iframe></div></figure><figure class="notion-asset-wrapper notion-asset-wrapper-embed notion-block-9c37c9f6ad3240dc9f3d631aa9529950"><div style="position:relative;display:flex;justify-content:center;align-self:center;width:100%;max-width:100%;flex-direction:column;height:320px"><iframe class="notion-asset-object-fit" src="https://platform.twitter.com/embed/Tweet.html?id=2082199414656127010" title="iframe embed" frameBorder="0" allowfullscreen="" loading="lazy" scrolling="auto"></iframe></div></figure><h4 class="notion-h notion-h3 notion-h-indent-1 notion-block-f84018c5fc6d46e19026b28dec2cbf11" data-id="f84018c5fc6d46e19026b28dec2cbf11"><span><div id="f84018c5fc6d46e19026b28dec2cbf11" class="notion-header-anchor"></div><a class="notion-hash-link" href="#f84018c5fc6d46e19026b28dec2cbf11" title="2. Thibault Sottiaux @thsottiaux"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">2. Thibault Sottiaux @thsottiaux</span></span></h4><div class="notion-text notion-block-46ea8fdd70434711be24faa1b81e3922">Thibault Sottiaux 今天是最重要的信息源。他先用一句“reset button”调侃额度重置，随后正式说明：所有 ChatGPT Work 和 Codex 用户的 usage limits 已重置；Sol 的典型使用时长预计会提升约 18%；此前因为调查临时暂停的五小时限制也会恢复。重点在于他解释了原因：GPT-5.6 Sol 更愿意长时间工作、调用更多工具、协调复杂 workflow 和 subagents，code mode 也带来了更多响应、更多缓存输入 token 和更高用量。另一条则是发布安全漏洞扫描、验证和修复相关的 CLI 和 TypeScript SDK，说明 Codex 周边的工程工具也在补齐。</div><figure class="notion-asset-wrapper notion-asset-wrapper-embed notion-block-d71fdd637c24417485fad98563fe1383"><div style="position:relative;display:flex;justify-content:center;align-self:center;width:100%;max-width:100%;flex-direction:column;height:320px"><iframe class="notion-asset-object-fit" src="https://platform.twitter.com/embed/Tweet.html?id=2082326593532473523" title="iframe embed" frameBorder="0" allowfullscreen="" loading="lazy" scrolling="auto"></iframe></div></figure><figure class="notion-asset-wrapper notion-asset-wrapper-embed notion-block-279d2c473b8b452d80c44397a4253b01"><div style="position:relative;display:flex;justify-content:center;align-self:center;width:100%;max-width:100%;flex-direction:column;height:320px"><iframe class="notion-asset-object-fit" src="https://platform.twitter.com/embed/Tweet.html?id=2082317452755751098" title="iframe embed" frameBorder="0" allowfullscreen="" loading="lazy" scrolling="auto"></iframe></div></figure><figure class="notion-asset-wrapper notion-asset-wrapper-embed notion-block-fb1336000581443daee4cb5e4d3ed9c7"><div style="position:relative;display:flex;justify-content:center;align-self:center;width:100%;max-width:100%;flex-direction:column;height:320px"><iframe class="notion-asset-object-fit" src="https://platform.twitter.com/embed/Tweet.html?id=2082241164850364555" title="iframe embed" frameBorder="0" allowfullscreen="" loading="lazy" scrolling="auto"></iframe></div></figure><h4 class="notion-h notion-h3 notion-h-indent-1 notion-block-40fa8f8841cd4cf399bbf8979111e78f" data-id="40fa8f8841cd4cf399bbf8979111e78f"><span><div id="40fa8f8841cd4cf399bbf8979111e78f" class="notion-header-anchor"></div><a class="notion-hash-link" href="#40fa8f8841cd4cf399bbf8979111e78f" title="3. Peter Yang @petergyang"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">3. Peter Yang @petergyang</span></span></h4><div class="notion-text notion-block-69287907624a47748f0a60f7cb895784">Peter Yang 今天一边讨论 Codex 的批评，一边发布自己的 Tastemaker 小产品。他说自己很喜欢 Codex，但也认可一些批评，并提出一个具体疑问：Sol Pro 是否比 Sol High 更好，为什么不在 Codex 里开放。另两条是他用 Claude Design 和 Claude Code 从粗略想法做出 Tastemaker 的过程：先写设计说明文件和 HTML spec，再用 Claude Design 做界面原型，最后用 Claude Code 构建和迭代。这是一个典型的 AI-native 产品构建流程案例。</div><figure class="notion-asset-wrapper notion-asset-wrapper-embed notion-block-55d7ebbb871d41dd80c8377bbfa56ad2"><div style="position:relative;display:flex;justify-content:center;align-self:center;width:100%;max-width:100%;flex-direction:column;height:320px"><iframe class="notion-asset-object-fit" src="https://platform.twitter.com/embed/Tweet.html?id=2082323512069685575" title="iframe embed" frameBorder="0" allowfullscreen="" loading="lazy" scrolling="auto"></iframe></div></figure><figure class="notion-asset-wrapper notion-asset-wrapper-embed notion-block-81310e8c77cf46e4b3d3094fa63bc9ea"><div style="position:relative;display:flex;justify-content:center;align-self:center;width:100%;max-width:100%;flex-direction:column;height:320px"><iframe class="notion-asset-object-fit" src="https://platform.twitter.com/embed/Tweet.html?id=2082254852600873376" title="iframe embed" frameBorder="0" allowfullscreen="" loading="lazy" scrolling="auto"></iframe></div></figure><figure class="notion-asset-wrapper notion-asset-wrapper-embed notion-block-f50b04b2a4d2456280a95de80fbb0e98"><div style="position:relative;display:flex;justify-content:center;align-self:center;width:100%;max-width:100%;flex-direction:column;height:320px"><iframe class="notion-asset-object-fit" src="https://platform.twitter.com/embed/Tweet.html?id=2082254840655405293" title="iframe embed" frameBorder="0" allowfullscreen="" loading="lazy" scrolling="auto"></iframe></div></figure><h4 class="notion-h notion-h3 notion-h-indent-1 notion-block-545c4bc64ccb46d79068e365a798f3eb" data-id="545c4bc64ccb46d79068e365a798f3eb"><span><div id="545c4bc64ccb46d79068e365a798f3eb" class="notion-header-anchor"></div><a class="notion-hash-link" href="#545c4bc64ccb46d79068e365a798f3eb" title="4. Madhu Guru @realmadhuguru"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">4. Madhu Guru @realmadhuguru</span></span></h4><div class="notion-text notion-block-e7c5dccaaa1040da927757a99134b3f0">Madhu Guru 的这条是半开玩笑，但很能说明现在工程师默认工作方式的变化：他看到一个 SWE 不用 Claude、不用 wispr、不用 tab key，只是“徒手写代码”。这不是产品发布，但它反过来说明 AI coding assistant 已经在一些团队里变成默认背景，纯手写代码反而像异常场景。</div><figure class="notion-asset-wrapper notion-asset-wrapper-embed notion-block-51c964d9609440cabc11d0ca0e379e37"><div style="position:relative;display:flex;justify-content:center;align-self:center;width:100%;max-width:100%;flex-direction:column;height:320px"><iframe class="notion-asset-object-fit" src="https://platform.twitter.com/embed/Tweet.html?id=2082112941814661236" title="iframe embed" frameBorder="0" allowfullscreen="" loading="lazy" scrolling="auto"></iframe></div></figure><h4 class="notion-h notion-h3 notion-h-indent-1 notion-block-9cb683f928104a3d97531d2104cd63a5" data-id="9cb683f928104a3d97531d2104cd63a5"><span><div id="9cb683f928104a3d97531d2104cd63a5" class="notion-header-anchor"></div><a class="notion-hash-link" href="#9cb683f928104a3d97531d2104cd63a5" title="5. Amjad Masad @amasad"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">5. Amjad Masad @amasad</span></span></h4><div class="notion-text notion-block-74c13acefb6340db972c9fc3bf4efc67">Amjad Masad 今天有三条。第一条是 rare books 相关，和 AI 主线关系不强；第二条更值得看，他把 SETI@Home 的捐算力搜索外星人类比到“捐算力搜索数学证明”，这延续了他昨天关于计算宇宙探索的想法；第三条只写了“1300 Elo”并附多个链接，JSON 里缺少上下文，按原始内容保留但不做额外推断。</div><figure class="notion-asset-wrapper notion-asset-wrapper-embed notion-block-0549e9dbde264a6399cf0d972b6b6495"><div style="position:relative;display:flex;justify-content:center;align-self:center;width:100%;max-width:100%;flex-direction:column;height:320px"><iframe class="notion-asset-object-fit" src="https://platform.twitter.com/embed/Tweet.html?id=2082317323445387514" title="iframe embed" frameBorder="0" allowfullscreen="" loading="lazy" scrolling="auto"></iframe></div></figure><figure class="notion-asset-wrapper notion-asset-wrapper-embed notion-block-fac6f5800f0c4b7cb26c9a40fe2325d0"><div style="position:relative;display:flex;justify-content:center;align-self:center;width:100%;max-width:100%;flex-direction:column;height:320px"><iframe class="notion-asset-object-fit" src="https://platform.twitter.com/embed/Tweet.html?id=2082316553740284060" title="iframe embed" frameBorder="0" allowfullscreen="" loading="lazy" scrolling="auto"></iframe></div></figure><figure class="notion-asset-wrapper notion-asset-wrapper-embed notion-block-33f669b522384ca3b4213547a304033f"><div style="position:relative;display:flex;justify-content:center;align-self:center;width:100%;max-width:100%;flex-direction:column;height:320px"><iframe class="notion-asset-object-fit" src="https://platform.twitter.com/embed/Tweet.html?id=2082316150273360316" title="iframe embed" frameBorder="0" allowfullscreen="" loading="lazy" scrolling="auto"></iframe></div></figure><h4 class="notion-h notion-h3 notion-h-indent-1 notion-block-6a60e5e53c8844678fe63763a75f2e32" data-id="6a60e5e53c8844678fe63763a75f2e32"><span><div id="6a60e5e53c8844678fe63763a75f2e32" class="notion-header-anchor"></div><a class="notion-hash-link" href="#6a60e5e53c8844678fe63763a75f2e32" title="6. Aaron Levie @levie"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">6. Aaron Levie @levie</span></span></h4><div class="notion-text notion-block-b42f9aec3135463ea5fed8de701e305a">Aaron Levie 今天两条都是 quote 形式。一条明确表达了对 Zuckerberg AI vision 的认可，另一条只有链接和转发内容，JSON 中没有更完整文本。结合他的 bio，至少能看出 Box 仍然把 AI 放在企业内容和业务工作流的核心位置，但这里不补外部信息，只保留原帖。</div><figure class="notion-asset-wrapper notion-asset-wrapper-embed notion-block-8c12ddb32a8a4f5ca389e7119e015f0b"><div style="position:relative;display:flex;justify-content:center;align-self:center;width:100%;max-width:100%;flex-direction:column;height:320px"><iframe class="notion-asset-object-fit" src="https://platform.twitter.com/embed/Tweet.html?id=2082168124733116537" title="iframe embed" frameBorder="0" allowfullscreen="" loading="lazy" scrolling="auto"></iframe></div></figure><figure class="notion-asset-wrapper notion-asset-wrapper-embed notion-block-35ad34f9f4d24ff98fe787f8543a3d18"><div style="position:relative;display:flex;justify-content:center;align-self:center;width:100%;max-width:100%;flex-direction:column;height:320px"><iframe class="notion-asset-object-fit" src="https://platform.twitter.com/embed/Tweet.html?id=2082114876873597239" title="iframe embed" frameBorder="0" allowfullscreen="" loading="lazy" scrolling="auto"></iframe></div></figure><h4 class="notion-h notion-h3 notion-h-indent-1 notion-block-1180a5ed078e4a798997469a02a5bbb3" data-id="1180a5ed078e4a798997469a02a5bbb3"><span><div id="1180a5ed078e4a798997469a02a5bbb3" class="notion-header-anchor"></div><a class="notion-hash-link" href="#1180a5ed078e4a798997469a02a5bbb3" title="7. Garry Tan @garrytan"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">7. Garry Tan @garrytan</span></span></h4><div class="notion-text notion-block-aaf358fc3f11401892e78cb9624181b1">Garry Tan 今天三条都不是 AI 主线。一条是关于生活态度：可以开玩笑，但面对人生、意图和世界观时要认真；另外两条是旧金山 Sanctuary City 政策相关。按“不筛选”的要求，这些内容保留，但不放进今天 AI 主线判断里。</div><figure class="notion-asset-wrapper notion-asset-wrapper-embed notion-block-9a98ae4e0b674ee4920269b6b45e9666"><div style="position:relative;display:flex;justify-content:center;align-self:center;width:100%;max-width:100%;flex-direction:column;height:320px"><iframe class="notion-asset-object-fit" src="https://platform.twitter.com/embed/Tweet.html?id=2082176112906711452" title="iframe embed" frameBorder="0" allowfullscreen="" loading="lazy" scrolling="auto"></iframe></div></figure><figure class="notion-asset-wrapper notion-asset-wrapper-embed notion-block-dd498fff08474abf8bea6b0e5ce7f37f"><div style="position:relative;display:flex;justify-content:center;align-self:center;width:100%;max-width:100%;flex-direction:column;height:320px"><iframe class="notion-asset-object-fit" src="https://platform.twitter.com/embed/Tweet.html?id=2082170934182756411" title="iframe embed" frameBorder="0" allowfullscreen="" loading="lazy" scrolling="auto"></iframe></div></figure><figure class="notion-asset-wrapper notion-asset-wrapper-embed notion-block-0d777683ae4c4c7494ea052ff6d65d84"><div style="position:relative;display:flex;justify-content:center;align-self:center;width:100%;max-width:100%;flex-direction:column;height:320px"><iframe class="notion-asset-object-fit" src="https://platform.twitter.com/embed/Tweet.html?id=2082169572212584877" title="iframe embed" frameBorder="0" allowfullscreen="" loading="lazy" scrolling="auto"></iframe></div></figure><h4 class="notion-h notion-h3 notion-h-indent-1 notion-block-6c8b36f626ae4e7ea28cf24921fbb620" data-id="6c8b36f626ae4e7ea28cf24921fbb620"><span><div id="6c8b36f626ae4e7ea28cf24921fbb620" class="notion-header-anchor"></div><a class="notion-hash-link" href="#6c8b36f626ae4e7ea28cf24921fbb620" title="8. Peter Steinberger @steipete"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">8. Peter Steinberger @steipete</span></span></h4><div class="notion-text notion-block-61889f757f80413aacb91c74539e18de">Peter Steinberger 今天只有一句判断：serving large models is hard。虽然短，但和今天 Codex 用量、Sol 效率、安全 SDK 这些内容放在一起很一致：大模型产品真正难的地方不只是模型聪明，而是如何稳定、低成本、安全地服务真实用户。</div><figure class="notion-asset-wrapper notion-asset-wrapper-embed notion-block-13e7e17a12534193b1c37097d41c1e14"><div style="position:relative;display:flex;justify-content:center;align-self:center;width:100%;max-width:100%;flex-direction:column;height:320px"><iframe class="notion-asset-object-fit" src="https://platform.twitter.com/embed/Tweet.html?id=2082337130299457652" title="iframe embed" frameBorder="0" allowfullscreen="" loading="lazy" scrolling="auto"></iframe></div></figure><h4 class="notion-h notion-h3 notion-h-indent-1 notion-block-45eb07ed20a44d8799784f6610941ae4" data-id="45eb07ed20a44d8799784f6610941ae4"><span><div id="45eb07ed20a44d8799784f6610941ae4" class="notion-header-anchor"></div><a class="notion-hash-link" href="#45eb07ed20a44d8799784f6610941ae4" title="9. Dan Shipper @danshipper"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">9. Dan Shipper @danshipper</span></span></h4><div class="notion-text notion-block-ddcecf15b05b4246b51e251c3a29dd35">Dan Shipper 今天最值得沉淀的是第三条：他在沙发上用 ChatGPT for Work 的 voice mode 完成了一篇 Codex 历史稿，从整理访谈、建立时间线，到写作、编辑、修改，全程没有碰键盘和鼠标。他说这对 creative work 非常强。另两条偏个人状态和轻量回应，和 AI 主线关系不强，但按全量要求保留。</div><figure class="notion-asset-wrapper notion-asset-wrapper-embed notion-block-a1d5125049684feca2ddbce27139cc92"><div style="position:relative;display:flex;justify-content:center;align-self:center;width:100%;max-width:100%;flex-direction:column;height:320px"><iframe class="notion-asset-object-fit" src="https://platform.twitter.com/embed/Tweet.html?id=2082273076352315440" title="iframe embed" frameBorder="0" allowfullscreen="" loading="lazy" scrolling="auto"></iframe></div></figure><figure class="notion-asset-wrapper notion-asset-wrapper-embed notion-block-fc077373b41b4f709bbbaf79e50c6b7e"><div style="position:relative;display:flex;justify-content:center;align-self:center;width:100%;max-width:100%;flex-direction:column;height:320px"><iframe class="notion-asset-object-fit" src="https://platform.twitter.com/embed/Tweet.html?id=2082270947793350785" title="iframe embed" frameBorder="0" allowfullscreen="" loading="lazy" scrolling="auto"></iframe></div></figure><figure class="notion-asset-wrapper notion-asset-wrapper-embed notion-block-28b4e14a160a41eab5223dc4db8c171c"><div style="position:relative;display:flex;justify-content:center;align-self:center;width:100%;max-width:100%;flex-direction:column;height:320px"><iframe class="notion-asset-object-fit" src="https://platform.twitter.com/embed/Tweet.html?id=2082130836485259530" title="iframe embed" frameBorder="0" allowfullscreen="" loading="lazy" scrolling="auto"></iframe></div></figure><h4 class="notion-h notion-h3 notion-h-indent-1 notion-block-5573268d6f9547c2a4572900a47ca876" data-id="5573268d6f9547c2a4572900a47ca876"><span><div id="5573268d6f9547c2a4572900a47ca876" class="notion-header-anchor"></div><a class="notion-hash-link" href="#5573268d6f9547c2a4572900a47ca876" title="10. Aditya Agarwal @adityaag"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">10. Aditya Agarwal @adityaag</span></span></h4><div class="notion-text notion-block-c51248c2450d4fc397b093a27c79f4cb">Aditya Agarwal 这条是个人反思。他回看十年前自己处在不确定中时得到的两个判断：想在世界上构建新东西，想和聪明、强烈的人一起工作。它不是 AI 产品动态，但作为 builder 的原始内容保留；放在今天的语境里，也和 AI-native IC / builder 市场的变化有轻微呼应。</div><figure class="notion-asset-wrapper notion-asset-wrapper-embed notion-block-447430faf83745d4ab915eae0dcc29f9"><div style="position:relative;display:flex;justify-content:center;align-self:center;width:100%;max-width:100%;flex-direction:column;height:320px"><iframe class="notion-asset-object-fit" src="https://platform.twitter.com/embed/Tweet.html?id=2082214798935326833" title="iframe embed" frameBorder="0" allowfullscreen="" loading="lazy" scrolling="auto"></iframe></div></figure><h3 class="notion-h notion-h2 notion-h-indent-0 notion-block-fd763fb2f5ee4d60b603e3bc7aa5574a" data-id="fd763fb2f5ee4d60b603e3bc7aa5574a"><span><div id="fd763fb2f5ee4d60b603e3bc7aa5574a" class="notion-header-anchor"></div><a class="notion-hash-link" href="#fd763fb2f5ee4d60b603e3bc7aa5574a" title="Blog 全量记录"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">Blog 全量记录</span></span></h3><div class="notion-text notion-block-e34612e011414c79a11b1417f4d44d9c">今天 JSON 中没有新增 blog。</div><h3 class="notion-h notion-h2 notion-h-indent-0 notion-block-0c1df58d16a24ae998578a5434e32a26" data-id="0c1df58d16a24ae998578a5434e32a26"><span><div id="0c1df58d16a24ae998578a5434e32a26" class="notion-header-anchor"></div><a class="notion-hash-link" href="#0c1df58d16a24ae998578a5434e32a26" title="Podcast 全量记录"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">Podcast 全量记录</span></span></h3><div class="notion-text notion-block-3dbfa36f49fd419caed29acef018db5a">今天 JSON 中没有新增 podcast。</div><h3 class="notion-h notion-h2 notion-h-indent-0 notion-block-1a77d2b6c5bd4d16b582f863c4eea690" data-id="1a77d2b6c5bd4d16b582f863c4eea690"><span><div id="1a77d2b6c5bd4d16b582f863c4eea690" class="notion-header-anchor"></div><a class="notion-hash-link" href="#1a77d2b6c5bd4d16b582f863c4eea690" title="今日沉淀结论"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">今日沉淀结论</span></span></h3><div class="notion-text notion-block-c40f6033b9164a91b2d9790c862a470a">今天的主线不是某个单点发布，而是 AI 工作流产品进入真实用量压力测试。Sol 这类更强的模型会做更多事，也会带来更复杂的额度、效率和透明度问题；会管理 agent、会把 voice mode / Claude Design / Claude Code 放进真实流程的人，正在变成更稀缺的工作能力。接下来值得继续盯的是：Codex 用量策略是否更清晰，Sol Pro / Sol High 是否在产品里分层，以及 AI-native IC 的工作方式会不会变成新的团队默认配置。</div><div class="notion-text notion-block-d97ac7ff3d104bc68891b9137e868652">Generated through the Follow Builders skill: <a target="_blank" rel="noopener noreferrer" class="notion-link" href="https://github.com/zarazhangrui/follow-builders">https://github.com/zarazhangrui/follow-builders</a></div></main></div>]]></content:encoded>
        </item>
        <item>
            <title><![CDATA[StateAct 深度分析 - 2026-07-29]]></title>
            <link>http://easyai.fyi/article/stateact-deep-analysis-2026-07-29</link>
            <guid>http://easyai.fyi/article/stateact-deep-analysis-2026-07-29</guid>
            <pubDate>Wed, 29 Jul 2026 00:00:00 GMT</pubDate>
            <description><![CDATA[StateAct 让电脑操作 Agent 优先读写程序状态，把屏幕点击降为必要时的补充。]]></description>
            <content:encoded><![CDATA[<div id="notion-article" class="mx-auto overflow-hidden "><main class="notion light-mode notion-page notion-block-3ac5eac3813c81f5b73bc837b9a9f692"><div class="notion-viewport"></div><div class="notion-collection-page-properties"></div><blockquote class="notion-quote notion-block-06592f662a844e529ff334724a3bdf72"><div><b>一句话判断</b>：StateAct 最有价值的地方，不是“用代码代替点击”这么简单，而是让 Agent 的执行、验收和长程状态都对准同一个对象：文件、数据库、DOM 和应用后端。它明显减少了屏幕操作带来的误差和成本，但最终成功率仍只有 26.9%；看得更准以后，推理错误成了主要障碍。</div></blockquote><h3 class="notion-h notion-h2 notion-h-indent-0 notion-block-6a2e7c92d9204bdca202029573fba07f" data-id="6a2e7c92d9204bdca202029573fba07f"><span><div id="6a2e7c92d9204bdca202029573fba07f" class="notion-header-anchor"></div><a class="notion-hash-link" href="#6a2e7c92d9204bdca202029573fba07f" title="原始材料"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">原始材料</span></span></h3><ul class="notion-list notion-list-disc notion-block-15e9660816bd41eeac68cd017634c376"><li><a target="_blank" rel="noopener noreferrer" class="notion-link" href="https://arxiv.org/abs/2607.22798">论文页面：arXiv:2607.22798</a></li></ul><ul class="notion-list notion-list-disc notion-block-c53a38f111d045b6816fdb66556b7294"><li><a target="_blank" rel="noopener noreferrer" class="notion-link" href="https://arxiv.org/html/2607.22798">论文 HTML 全文</a></li></ul><ul class="notion-list notion-list-disc notion-block-358eab6952b149f1a542809ccbf0f4d3"><li><a target="_blank" rel="noopener noreferrer" class="notion-link" href="https://arxiv.org/pdf/2607.22798">论文 PDF</a></li></ul><ul class="notion-list notion-list-disc notion-block-4f5ee0590ac04d549637c94d7e7cdd5d"><li><a target="_blank" rel="noopener noreferrer" class="notion-link" href="https://huggingface.co/papers/2607.22798">Hugging Face 论文页</a></li></ul><ul class="notion-list notion-list-disc notion-block-72d7f9d376f74f1ca622c1c20f3ec1f3"><li><a target="_blank" rel="noopener noreferrer" class="notion-link" href="https://huggingface.co/papers/date/2026-07-28">2026 年 7 月 28 日 Hugging Face Daily Papers</a></li></ul><div class="notion-text notion-block-e6dddfe416574d82a8318dfdc0dfb088">论文来自 Salesforce AI Research，2026 年 7 月 24 日提交。到 7 月 29 日，我在 arXiv 页面、论文正文和 Hugging Face 条目中没有看到官方代码入口，因此下面的实验结果均按作者报告理解，不写成独立复现结果。</div><h3 class="notion-h notion-h2 notion-h-indent-0 notion-block-fcececdd667248069467c90ffff59215" data-id="fcececdd667248069467c90ffff59215"><span><div id="fcececdd667248069467c90ffff59215" class="notion-header-anchor"></div><a class="notion-hash-link" href="#fcececdd667248069467c90ffff59215" title="为什么今天选它"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">为什么今天选它</span></span></h3><div class="notion-text notion-block-fe872301eed8446f98c38c964a4cb3c0">7 月 28 日的 Agent 候选里，JarvisHub 把画布当作创作 Agent 的共享状态，Multi-Agent Protocol Distillation 研究搜索轨迹蒸馏，Interactive Reward Agent 则用环境状态评估 GUI 任务。StateAct 更值得今天拆，因为它同时碰到了 computer use、harness、状态管理、工具路由、独立验收和长上下文，而且同模型对照、拆解实验与失败审计都比较完整。</div><div class="notion-text notion-block-22bcc41d8fa8460d8ee337c5ee4e6b9d">它也和最近几次沉淀错开了。前几天研究的是世界模型、上下文记忆和训练时 skill；StateAct 讨论的是运行中的 Agent 到底该观察什么、改什么、凭什么宣布完成。这个问题直接影响真实产品的可靠性与成本。</div><h3 class="notion-h notion-h2 notion-h-indent-0 notion-block-0e12fdda76fa48f48e9e0f6e3d20c89b" data-id="0e12fdda76fa48f48e9e0f6e3d20c89b"><span><div id="0e12fdda76fa48f48e9e0f6e3d20c89b" class="notion-header-anchor"></div><a class="notion-hash-link" href="#0e12fdda76fa48f48e9e0f6e3d20c89b" title="它在解决什么问题"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">它在解决什么问题</span></span></h3><div class="notion-text notion-block-9b49c73cbf504ffa851940356cd2aeb8">多数 computer-use Agent 的基本循环是：看截图，判断下一步，移动鼠标或输入文字，再看一张新截图。这种方式把屏幕当成事实来源，可屏幕只是程序状态的一次渲染。</div><div class="notion-text notion-block-4cfb441cf26e4449a236b8e2c66c37f8">同一个表格总计显示为 1,240，单元格里可能是公式，也可能是手填数字；隐藏行、完整精度和未保存内容也可能不在画面中。对一步任务，这种损失未必致命。连续做上百步时，每次都从不完整画面猜测，误差会不断积累。更麻烦的是，任务最终交付的是保存后的文件或应用状态，不是最后一张截图。<a target="_blank" rel="noopener noreferrer" class="notion-link" href="https://arxiv.org/html/2607.22798#S3">论文第 3 节</a>把这一区别称为 state-grounding，也就是“对程序状态落地”。</div><figure class="notion-asset-wrapper notion-asset-wrapper-image notion-block-038356e4ec494ca5a3009c4317a32499"><div style="position:relative;display:flex;justify-content:center;align-self:center;width:100%;max-width:100%;flex-direction:column"><img style="object-fit:cover" src="https://arxiv.org/html/2607.22798v1/x3.png?t=038356e4-ec49-4ca5-a300-9c4317a32499" alt="notion image" loading="lazy" decoding="async"/></div></figure><div class="notion-text notion-block-5fd34e3f6ca142b495eb7ae9691ab595"><em>图 1：论文 Figure 3。左侧通过截图与坐标操作表格，看到的是渲染结果；右侧直接读取工作簿中的公式并保存，拿到的是交付物本身的状态。来源：</em><em><a target="_blank" rel="noopener noreferrer" class="notion-link" href="https://arxiv.org/html/2607.22798#S3">论文 HTML</a></em><em>。</em></div><div class="notion-text notion-block-34c6ec8df712487b94da79bbd98d6bca">StateAct 的判断很明确：只要目标能通过文件、数据库、DOM 或应用后端表达，就优先走状态通道；确实只能靠视觉完成的步骤，再交给 GUI Agent。</div><h3 class="notion-h notion-h2 notion-h-indent-0 notion-block-96296df91b2c450b955f737a5aaca9f5" data-id="96296df91b2c450b955f737a5aaca9f5"><span><div id="96296df91b2c450b955f737a5aaca9f5" class="notion-header-anchor"></div><a class="notion-hash-link" href="#96296df91b2c450b955f737a5aaca9f5" title="系统如何运作"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">系统如何运作</span></span></h3><h4 class="notion-h notion-h3 notion-h-indent-1 notion-block-4c6b0c3166f741d18cab5286cabd1f98" data-id="4c6b0c3166f741d18cab5286cabd1f98"><span><div id="4c6b0c3166f741d18cab5286cabd1f98" class="notion-header-anchor"></div><a class="notion-hash-link" href="#4c6b0c3166f741d18cab5286cabd1f98" title="1. 主 Agent 先找状态，再动手"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">1. 主 Agent 先找状态，再动手</span></span></h4><div class="notion-text notion-block-bfee5a5db1894459a712bfde2caf3efe">主 Agent 没有鼠标和键盘控制权。它能用持久 shell、文件编辑器、Python、只读图片查看、计划清单、结束动作和委派工具。面对一个应用，它先判断状态存在哪里；不确定时，就用 <code class="notion-inline-code">find</code>、<code class="notion-inline-code">grep</code>、<code class="notion-inline-code">sqlite3</code> 一类手段探查文件和数据库。<a target="_blank" rel="noopener noreferrer" class="notion-link" href="https://arxiv.org/html/2607.22798#S4.SS1">论文第 4.1 节</a>特别说明，探查只能定位状态，不能凭空告诉 Agent 正确答案。</div><div class="notion-text notion-block-52f3bd55ec2940718f1c42fcf91c22ea">这个约束改变了默认路线。普通 GUI Agent 往往先尝试点击，失败后才找别的接口；StateAct 先操作真实交付物，屏幕是补充手段。</div><h4 class="notion-h notion-h3 notion-h-indent-1 notion-block-225dfd13926e4653ba5ae2589c1ec4a1" data-id="225dfd13926e4653ba5ae2589c1ec4a1"><span><div id="225dfd13926e4653ba5ae2589c1ec4a1" class="notion-header-anchor"></div><a class="notion-hash-link" href="#225dfd13926e4653ba5ae2589c1ec4a1" title="2. GUI 和浏览器 Agent 是专门工种"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">2. GUI 和浏览器 Agent 是专门工种</span></span></h4><div class="notion-text notion-block-0abad159e47e4eb09b7106eab99655ec">遇到拖拽、不可脚本化的弹窗、只存在于画面上的信息，主 Agent 才委派 GUI 子 Agent。论文的 108 个 OSWorld 2.0 任务里，只有 28 个任务调用过它；按主 Agent 步数算，GUI 调用占 1.1%。把子 Agent 内部步骤也算进去，视觉操作约占全部模型轮次的 11%。</div><div class="notion-text notion-block-fc82bebb1f7240209b8fc875a0f46759">网页任务另有浏览器 Agent，直接读 DOM、运行 JavaScript、按 CSS 选择器操作。它仍然遵守“先用结构化状态”的原则，并不是换一个 Agent 继续看截图。</div><figure class="notion-asset-wrapper notion-asset-wrapper-image notion-block-80ea37c4d0de4fa0b0940a7271102ace"><div style="position:relative;display:flex;justify-content:center;align-self:center;width:100%;max-width:100%;flex-direction:column"><img style="object-fit:cover" src="https://arxiv.org/html/2607.22798v1/x2.png?t=80ea37c4-d0de-4fa0-b094-0a7271102ace" alt="notion image" loading="lazy" decoding="async"/></div></figure><div class="notion-text notion-block-d378d14897474381816bec53d201f0f4"><em>图 2：论文 Figure 2。主 Agent 负责程序状态，GUI、网页和通用子 Agent 只接收聚焦后的子任务；完成前还要经过独立验收。来源：</em><em><a target="_blank" rel="noopener noreferrer" class="notion-link" href="https://arxiv.org/html/2607.22798v1/x2.png">论文架构图</a></em><em>。</em></div><h4 class="notion-h notion-h3 notion-h-indent-1 notion-block-95e5519c70f841d8aca2795ec7853edc" data-id="95e5519c70f841d8aca2795ec7853edc"><span><div id="95e5519c70f841d8aca2795ec7853edc" class="notion-header-anchor"></div><a class="notion-hash-link" href="#95e5519c70f841d8aca2795ec7853edc" title="3. 完成不是一句“做完了”，而是重新检查交付物"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">3. 完成不是一句“做完了”，而是重新检查交付物</span></span></h4><div class="notion-text notion-block-507c457abd864977be00444c2648204d">主 Agent 调用 finish 后，会启动一个只读的 finish gate。它只看到原始任务和当前机器状态，看不到主 Agent 的聊天记录、计划和完成理由，也不能修改文件。它必须自己找到任务真正要求的交付物，再检查文件是否存在、路径是否正确、内容是否保存、格式是否符合要求。</div><div class="notion-text notion-block-40f9b1c31d904564945b8e7286f40fed">这套验收有两个值得借鉴的限制。第一，不接受 Agent 自己额外写出的“证明文件”，因为那可能只是自证循环；证据必须来自任务指定的真实交付物。第二，验收失败后最多重做三轮，避免无限反思。<a target="_blank" rel="noopener noreferrer" class="notion-link" href="https://arxiv.org/html/2607.22798#S4.SS2">论文第 4.2 节</a>称它为 narration-blind，也就是不听执行者讲故事，只看结果。</div><figure class="notion-asset-wrapper notion-asset-wrapper-image notion-block-40a783fa5540483393d825ab43e2f488"><div style="position:relative;display:flex;justify-content:center;align-self:center;width:100%;max-width:100%;flex-direction:column"><img style="object-fit:cover" src="https://arxiv.org/html/2607.22798v1/x10.png?t=40a783fa-5540-4833-93d8-25ab43e2f488" alt="notion image" loading="lazy" decoding="async"/></div></figure><div class="notion-text notion-block-9f8a9e03c8cb47ac97d571b3492badca"><em>图 3：论文 Figure 4。Agent 声称已导入 14 个日历事件，finish gate 读取实际存储后发现只有 8 个，于是退回重做。它擅长抓住漏存、错路径和格式错误。来源：</em><em><a target="_blank" rel="noopener noreferrer" class="notion-link" href="https://arxiv.org/html/2607.22798v1/x10.png">论文 Figure 4</a></em><em>。</em></div><h4 class="notion-h notion-h3 notion-h-indent-1 notion-block-d32df03362424abb8c2fcf12b0e6dc2e" data-id="d32df03362424abb8c2fcf12b0e6dc2e"><span><div id="d32df03362424abb8c2fcf12b0e6dc2e" class="notion-header-anchor"></div><a class="notion-hash-link" href="#d32df03362424abb8c2fcf12b0e6dc2e" title="4. 长程状态靠三件朴素的东西维持"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">4. 长程状态靠三件朴素的东西维持</span></span></h4><div class="notion-text notion-block-f7de89f2a3dc450b9face80efc648a09">每个任务最多允许主 Agent 运行 200 轮，平均还会委派 4.2 次子任务。StateAct 没有把所有轨迹一直留在主上下文里，而是让每个子 Agent 使用新的上下文，只返回简短结果；接近上下文上限时压缩较早内容并移除旧图片；计划清单单独保存，每轮重新注入。<a target="_blank" rel="noopener noreferrer" class="notion-link" href="https://arxiv.org/html/2607.22798#S4.SS3">论文第 4.3 节</a>给出的这三项机制并不新奇，却把“任务事实”和“过程噪声”分开了。</div><h3 class="notion-h notion-h2 notion-h-indent-0 notion-block-03b1a40274ef49e0b3ba993f38fb85d4" data-id="03b1a40274ef49e0b3ba993f38fb85d4"><span><div id="03b1a40274ef49e0b3ba993f38fb85d4" class="notion-header-anchor"></div><a class="notion-hash-link" href="#03b1a40274ef49e0b3ba993f38fb85d4" title="实验结果应该怎么看"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">实验结果应该怎么看</span></span></h3><div class="notion-text notion-block-fe56ec47f2b14cbc81397ee7ffec7a97">论文在 OSWorld 2.0 的 108 个长程桌面任务上评测。使用同一个 Claude Opus 4.8 时，参考 GUI harness 的完全成功率是 20.6%，StateAct 是 26.9%；部分成功分从 54.8% 提升到 61.6%。平均输出 token 从 224K 降到 100K，作者按当时价格估算的单任务成本从约 72 美元降到 7.8 美元，约为原来的九分之一。<a target="_blank" rel="noopener noreferrer" class="notion-link" href="https://arxiv.org/html/2607.22798#S5">论文实验部分</a>给出了同模型对照。</div><figure class="notion-asset-wrapper notion-asset-wrapper-image notion-block-a30aaa68b26f49b08b3c538c20da542a"><div style="position:relative;display:flex;justify-content:center;align-self:center;width:100%;max-width:100%;flex-direction:column"><img style="object-fit:cover" src="https://arxiv.org/html/2607.22798v1/x12.png?t=a30aaa68-b26f-49b0-8b3c-538c20da542a" alt="notion image" loading="lazy" decoding="async"/></div></figure><div class="notion-text notion-block-c136066cb07f418eab019597cdac5e89"><em>图 4：论文 Figure 6。横轴是单任务成本，纵轴分别是完全成功率和部分得分。StateAct 的点位同时向“更便宜”和“更准确”移动。所有数字来自作者汇总的公开轨迹与自家运行结果。来源：</em><em><a target="_blank" rel="noopener noreferrer" class="notion-link" href="https://arxiv.org/html/2607.22798v1/x12.png">论文 Figure 6</a></em><em>。</em></div><div class="notion-text notion-block-a44197f7f7cb41c6abb477606a03d66e">最能说明问题的不是总分，而是拆解实验：</div><ul class="notion-list notion-list-disc notion-block-f521ad11d07d491ab8c1cde96ccaad48"><li>拿掉 state-first 执行，部分分从 61.6% 降到 51.3%，甚至低于原始 GUI harness 的 54.8%；</li></ul><ul class="notion-list notion-list-disc notion-block-4c6d767a75be42a8a511a50fb20dc8d9"><li>拿掉 finish gate，部分分降到 57.5%；</li></ul><ul class="notion-list notion-list-disc notion-block-9311576a3e074295a9c9dc6ca8fd6c33"><li>拿掉计划与压缩，部分分降到 58.7%；</li></ul><ul class="notion-list notion-list-disc notion-block-7381d819fdd64a33859346c50665be5e"><li>只留 shell、不用 GUI 子 Agent、委派和验收，部分分只有 45.9%。</li></ul><div class="notion-text notion-block-681c90ba277f496fafe1944df0b9de81">这些结果说明，直接操作状态是最大贡献，但“全用代码”并不成立。视觉兜底、独立验收和长程管理合在一起才有效。另一个有意思的结果是，扁平委派的部分分为 61.6%，两种递归版本只有 57.6% 和 57.9%；递归分支仅在 108 个任务中的 7 个被触发，而且没有真正嵌套。论文据此支持“这套任务先保持扁平”，没有把它夸成多层 Agent 永远无用。</div><div class="notion-text notion-block-84fca6281f314d8dae0e98391250d409">短任务上的优势也很小。OSWorld-Verified 中，StateAct 与参考系统的完全成功率是 78.4% 对 77.3%。这和论文的主张相符：状态通道主要在长链任务中减少累积误差，不是任何 GUI 任务都能吃到同样收益。</div><h3 class="notion-h notion-h2 notion-h-indent-0 notion-block-4dc48d76fa8d487f9008bc0ea493ad8f" data-id="4dc48d76fa8d487f9008bc0ea493ad8f"><span><div id="4dc48d76fa8d487f9008bc0ea493ad8f" class="notion-header-anchor"></div><a class="notion-hash-link" href="#4dc48d76fa8d487f9008bc0ea493ad8f" title="真正的新意在哪里"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">真正的新意在哪里</span></span></h3><div class="notion-text notion-block-7bbc3c7c45a8471daa25bc2d1462bdc8">用代码作为 Agent 动作、把 API 和 GUI 混合、增加独立检查器、压缩上下文、委派新子 Agent，这些零件以前都有。StateAct 没有发明它们。</div><div class="notion-text notion-block-a83215d2c78042b9a431c4c497aa49ce">它真正做对的是统一工作对象和默认路由。执行时改真实状态，完成时重读真实状态，长程管理时保留状态事实；视觉操作只处理状态通道确实碰不到的部分。常见混合 Agent 往往把 GUI 与 API 当成两个并列工具，让模型每一步自由选择。StateAct 把状态通道放在主循环里，把 GUI 隔离成少量、明确的例外。</div><div class="notion-text notion-block-8b0348903cbb4ecda9d071e884c53dc5">finish gate 的价值也不在“多一个模型来反思”。它故意看不到执行历史，减少被主 Agent 叙述带偏的机会；同时要求重新定位真实交付物，避免 Agent 用自己制造的旁证骗过检查。这里的新意更像系统边界设计，而不是一种新推理算法。</div><h3 class="notion-h notion-h2 notion-h-indent-0 notion-block-2a45118b6f6f4334bc6f9e6d087098e1" data-id="2a45118b6f6f4334bc6f9e6d087098e1"><span><div id="2a45118b6f6f4334bc6f9e6d087098e1" class="notion-header-anchor"></div><a class="notion-hash-link" href="#2a45118b6f6f4334bc6f9e6d087098e1" title="对产品和研发的启发"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">对产品和研发的启发</span></span></h3><div class="notion-text notion-block-d471f2919aec42559d349c651a2c2370">做真实业务 Agent 时，可以先给每类操作定义“权威状态”是什么。表格任务看工作簿内容，日历任务看事件存储，网页任务看 DOM 与后端返回，只有版式和视觉效果才以渲染结果为准。Agent 不必在每一步都模拟人类操作界面。</div><div class="notion-text notion-block-e3f14dad4a624dfea73a78e339ae0333">完成条件也该写在交付物上。不要用“模型说完成”或“最后一张截图看起来对”作为验收；至少检查路径、保存状态、结构、数量和格式。能写确定性检查的地方就别交给同一个模型再猜一遍。</div><div class="notion-text notion-block-36d9cfdecb20408d9df2c8643bf8a342">执行者与验收者应隔离上下文。验收者只拿任务和结果，读权限够用就不要给写权限。即使它查不出所有错误，也能以较低成本拦住一批漏存、错路径和格式问题。</div><div class="notion-text notion-block-0a89640a769e49f7b519abf00860edfa">长任务的计划要放在消息历史之外。子任务用短上下文执行，返回结论和证据；图片、试错记录和临时日志不必长期占据主 Agent 的注意力。这里的重点不是压缩得多聪明，而是别把计划和过程垃圾混在一起。</div><h3 class="notion-h notion-h2 notion-h-indent-0 notion-block-d0a1e0582bf64767ace2e19cc11c9b64" data-id="d0a1e0582bf64767ace2e19cc11c9b64"><span><div id="d0a1e0582bf64767ace2e19cc11c9b64" class="notion-header-anchor"></div><a class="notion-hash-link" href="#d0a1e0582bf64767ace2e19cc11c9b64" title="风险、局限和尚未验证的问题"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">风险、局限和尚未验证的问题</span></span></h3><div class="notion-text notion-block-6aedd980be0f4b8997a539f597e63fc1">StateAct 把准确性墙从感知推向了推理，却没有穿过去。108 个任务中只有 29 个完全成功，79 个仍不完美。finish gate 对 76 个已进入验收但最终不完美的任务，只正确拒绝了 8 个，却放过了 68 个，漏过率约 90%。原因不是它相信了主 Agent 的叙述，而是两者可能用同一种方式理解错了原始材料。没有标准答案时，通用检查器擅长结构，难以判断数值是否真的正确。</div><figure class="notion-asset-wrapper notion-asset-wrapper-image notion-block-b90a314f48e140099a9938d3ba7f0cfa"><div style="position:relative;display:flex;justify-content:center;align-self:center;width:100%;max-width:100%;flex-direction:column"><img style="object-fit:cover" src="https://arxiv.org/html/2607.22798v1/x14.png?t=b90a314f-48e1-4009-9a99-38d3ba7f0cfa" alt="notion image" loading="lazy" decoding="async"/></div></figure><div class="notion-text notion-block-83659f0de55947f5b34cdab250340242"><em>图 5：论文 Figure 8。79 个未完全成功任务中，38 个主要失败原因是推理错误；另有验收薄弱、视觉能力不足和指令歧义。分类来自作者人工审计，不是自动标注。来源：</em><em><a target="_blank" rel="noopener noreferrer" class="notion-link" href="https://arxiv.org/html/2607.22798v1/x14.png">论文 Figure 8</a></em><em>。</em></div><div class="notion-text notion-block-10fa7290bf8e4ebd91ed1f21985a2613">状态通道也有清楚的边界。图像编辑、版式、CAD、拖拽和只存在于画面的信息，本来就需要视觉判断。论文中 human-in-the-loop 类任务的完全成功率仍是 0%，多模态编辑也明显弱于状态可查询的任务。把 GUI 降为例外，不等于视觉能力可以不要。</div><div class="notion-text notion-block-c271be48f39c4022ac1c568235142d4b">直接读写文件、数据库和应用后端会扩大权限与安全风险。一个被诱导的 Agent 不只是点错按钮，还可能批量修改真实数据。论文主要评估完成率，没有系统研究权限隔离、审计、恶意输入、回滚和并发写入。</div><div class="notion-text notion-block-a015382f3a1345b4ac6d66a7d0f7bcfb">证据范围也要收紧。主结果依赖 OSWorld 2.0 的 108 个任务和 Claude Opus 4.8；较弱的 Sonnet 4.6 上，部分分只从 41.5% 变成 42.0%。成本数字还会随模型定价与缓存策略变化。加上官方代码尚未公开，目前能确认的是论文内部对照完整，不能确认第三方能否按同样成本复现。</div><h3 class="notion-h notion-h2 notion-h-indent-0 notion-block-eeb7b7ffbf5f4135a65f43e4f1582823" data-id="eeb7b7ffbf5f4135a65f43e4f1582823"><span><div id="eeb7b7ffbf5f4135a65f43e4f1582823" class="notion-header-anchor"></div><a class="notion-hash-link" href="#eeb7b7ffbf5f4135a65f43e4f1582823" title="今日沉淀"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">今日沉淀</span></span></h3><ol start="1" class="notion-list notion-list-numbered notion-block-c4356e180dae40fc893b9c70a8fb7359"><li>先找任务的权威状态，再决定是否需要看屏幕。</li></ol><ol start="2" class="notion-list notion-list-numbered notion-block-520ce43349d342308504509d3189ab89"><li>GUI 应该是处理视觉例外的专门工具，不必占据主循环。</li></ol><ol start="3" class="notion-list notion-list-numbered notion-block-77eb55373d534d7faf6b6ac2fd0955f6"><li>验收者只看任务和真实交付物，不看执行者的自述。</li></ol><ol start="4" class="notion-list notion-list-numbered notion-block-b8a1b2c9975d4d1da53cafd9eef9256b"><li>通用 finish gate 能查结构错误，查不了共同理解造成的数值错误。</li></ol><ol start="5" class="notion-list notion-list-numbered notion-block-bc17591db4834ebe8d2b4e2844256d9e"><li>长任务优先外置计划、隔离子任务上下文，再考虑更深的多 Agent 层级。</li></ol></main></div>]]></content:encoded>
        </item>
        <item>
            <title><![CDATA[Agent 成本、安全和工作流成主线 - 2026-07-28]]></title>
            <link>http://easyai.fyi/article/follow-builders-ai-summary-2026-07-28</link>
            <guid>http://easyai.fyi/article/follow-builders-ai-summary-2026-07-28</guid>
            <pubDate>Tue, 28 Jul 2026 00:00:00 GMT</pubDate>
            <description><![CDATA[本轮 follow-builders 新增 14 个 X builders、29 条 tweet、0 篇 blog 和 1 期 podcast，已按输出全量沉淀。今天主线是 agent 从演示进入真实工作流，同时成本口径、安全边界、模型权重和 AI-native 工作界面成为讨论重点。]]></description>
            <content:encoded><![CDATA[<div id="notion-article" class="mx-auto overflow-hidden "><main class="notion light-mode notion-page notion-block-3ab5eac3813c81b1801cec39d57a1dca"><div class="notion-viewport"></div><div class="notion-collection-page-properties"></div><h3 class="notion-h notion-h2 notion-h-indent-0 notion-block-9ac789bc89924827b6627cf540f263c7" data-id="9ac789bc89924827b6627cf540f263c7"><span><div id="9ac789bc89924827b6627cf540f263c7" class="notion-header-anchor"></div><a class="notion-hash-link" href="#9ac789bc89924827b6627cf540f263c7" title="今日主线"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">今日主线</span></span></h3><div class="notion-text notion-block-0b0040b98ab94340a996a7b2849e1c90">今天这批内容的主线很清楚：agent 正在从“能做 demo”进入真实工作流。Codex 已经被拿来远程改视频、循环看 Slack 反馈、做版本迭代；Claude Code 被拿去做旅行复盘；不同 agent 之间甚至开始出现“一个报 bug，另一个修 bug”的协作。与此同时，大家开始关心更底层的问题：成本要按任务算，agent 要跑在更强隔离里，模型权重和 eval 透明度会影响信任。</div><h3 class="notion-h notion-h2 notion-h-indent-0 notion-block-d04e9641c0994adda79524314118c2d1" data-id="d04e9641c0994adda79524314118c2d1"><span><div id="d04e9641c0994adda79524314118c2d1" class="notion-header-anchor"></div><a class="notion-hash-link" href="#d04e9641c0994adda79524314118c2d1" title="重点解读"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">重点解读</span></span></h3><div class="notion-text notion-block-101117fc6d624503be86ea7819174389">第一条线是 Codex / ChatGPT Work 的采用速度和使用额度。Thibault Sottiaux 说 Codex 和 ChatGPT Work 的付费用户额度已重置，并把这和 ChatGPT Work 的快速采用放在一起。Peter Yang 转述 OpenAI DevEx Jason 的例子更具体：人在骑车时，用手机远程让 Codex 通过 computer use 改 launch video、导出、发回 Slack，还每 30 分钟检查反馈并生成后续版本。这类案例比“会写代码”更重要，因为它说明 agent 正在进入多人协作和交付闭环。</div><div class="notion-text notion-block-04a4a6f768fc4665852b447b037878b9">第二条线是成本和评测口径。Swyx 直接说 $/input-output token 已经过时，应该看 $/task。Guillermo Rauch 也从 Vercel benchmark 角度谈 cybersecurity model 的价格性能，认为 Grok 4.5 在 price-performance 上很强，同时 Sol 仍是 frontier。这些讨论都指向同一个变化：模型比较不再只看单次 token 单价，也要看它完成一个真实任务花多少钱、稳不稳、能不能被安全运行。</div><div class="notion-text notion-block-ccc60f174fff46d5be5993da75de7c48">第三条线是 agent 安全边界。Guillermo Rauch 提到 Kimi paper 的信号：container-level isolation 不够，实验里 agent 能把底层机器打崩到 kernel panic；他认为 Firecracker microVM 这类边界更安全。Peter Steinberger 也提到团队做了大量安全工作，但外界仍担心“不安全”。这说明 agent 越接近真实系统，安全讨论越会从抽象原则变成运行时隔离、权限、审计和恢复能力。</div><div class="notion-text notion-block-837dd4e031c1454d9c8d69fdcb37382e">第四条线是 AI-native 产品形态。AI &amp; I 这期里，Granola CEO 的核心判断是 meeting notes 不是终点，真正的问题是未来工作的界面会变成什么。Dan Shipper 提到工作会分成两类：一种是在 Slack 里公开、多人、异步地派活；另一种是在 Codex 或 Claude desktop 这类同屏环境里，人和 agent 深度协作。这个框架很适合解释今天的 X 动态：大家不是只在找新工具，而是在找新的工作界面。</div><h3 class="notion-h notion-h2 notion-h-indent-0 notion-block-f9cf81da725943e6a2fbe3451e2d6ad4" data-id="f9cf81da725943e6a2fbe3451e2d6ad4"><span><div id="f9cf81da725943e6a2fbe3451e2d6ad4" class="notion-header-anchor"></div><a class="notion-hash-link" href="#f9cf81da725943e6a2fbe3451e2d6ad4" title="X Builders 全量记录"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">X Builders 全量记录</span></span></h3><h4 class="notion-h notion-h3 notion-h-indent-1 notion-block-21df71d96c5a40a088a3f171fa6d51ff" data-id="21df71d96c5a40a088a3f171fa6d51ff"><span><div id="21df71d96c5a40a088a3f171fa6d51ff" class="notion-header-anchor"></div><a class="notion-hash-link" href="#21df71d96c5a40a088a3f171fa6d51ff" title="1. Swyx @swyx"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">1. Swyx @swyx</span></span></h4><div class="notion-text notion-block-dfd64a844b794844844352c1efad0b22">Swyx 今天三条都围绕模型评估和 agent 生态。一条是对 Kimi / Moonshot 的轻量回应；最有价值的是他强调模型成本口径应该从 $/token 转向 $/task，因为真实产品买的是任务完成，不是 token；另一条则回看自己的 agent lab thesis，认为 eval、routing、interactivity、ROI 这些方向判断对了，但 Claude Code “意外开源”后并没有明显改变竞品路线，也说明 agent 产品的护城河不只在代码本身。</div><figure class="notion-asset-wrapper notion-asset-wrapper-embed notion-block-289e7bc092dc4da7ae86716b7d87a9d8"><div style="position:relative;display:flex;justify-content:center;align-self:center;width:100%;max-width:100%;flex-direction:column;height:320px"><iframe class="notion-asset-object-fit" src="https://platform.twitter.com/embed/Tweet.html?id=2081979163117052311" title="iframe embed" frameBorder="0" allowfullscreen="" loading="lazy" scrolling="auto"></iframe></div></figure><figure class="notion-asset-wrapper notion-asset-wrapper-embed notion-block-d27aa8994fbe4facb6c3328c629fa45a"><div style="position:relative;display:flex;justify-content:center;align-self:center;width:100%;max-width:100%;flex-direction:column;height:320px"><iframe class="notion-asset-object-fit" src="https://platform.twitter.com/embed/Tweet.html?id=2081904230768816487" title="iframe embed" frameBorder="0" allowfullscreen="" loading="lazy" scrolling="auto"></iframe></div></figure><figure class="notion-asset-wrapper notion-asset-wrapper-embed notion-block-9b51d004f91148ad930838ac6081bb55"><div style="position:relative;display:flex;justify-content:center;align-self:center;width:100%;max-width:100%;flex-direction:column;height:320px"><iframe class="notion-asset-object-fit" src="https://platform.twitter.com/embed/Tweet.html?id=2081890955070980416" title="iframe embed" frameBorder="0" allowfullscreen="" loading="lazy" scrolling="auto"></iframe></div></figure><h4 class="notion-h notion-h3 notion-h-indent-1 notion-block-afe4e266c4a14311b2a8743f16ef784d" data-id="afe4e266c4a14311b2a8743f16ef784d"><span><div id="afe4e266c4a14311b2a8743f16ef784d" class="notion-header-anchor"></div><a class="notion-hash-link" href="#afe4e266c4a14311b2a8743f16ef784d" title="2. Thibault Sottiaux @thsottiaux"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">2. Thibault Sottiaux @thsottiaux</span></span></h4><div class="notion-text notion-block-44a597151b7b4bbbad6274943fc1567a">Thibault 今天围绕 ChatGPT Work 和 Codex 的使用节奏。他先说在庆祝 ChatGPT Work 的快速采用，并暗示会重置额度；随后确认所有 Codex 和 ChatGPT Work 付费用户的 usage limits 已重置；最后说自己短暂离开 X 休息，后面还有更多 ChatGPT 和 Codex 动态。这些内容说明 OpenAI 正在很主动地运营这批工作流产品，不只是发布一次功能。</div><figure class="notion-asset-wrapper notion-asset-wrapper-embed notion-block-2ffc61ba54094735bfe1c31e25b800dc"><div style="position:relative;display:flex;justify-content:center;align-self:center;width:100%;max-width:100%;flex-direction:column;height:320px"><iframe class="notion-asset-object-fit" src="https://platform.twitter.com/embed/Tweet.html?id=2081979033261412537" title="iframe embed" frameBorder="0" allowfullscreen="" loading="lazy" scrolling="auto"></iframe></div></figure><figure class="notion-asset-wrapper notion-asset-wrapper-embed notion-block-8761c69e4dbd4f36883739af91326c8c"><div style="position:relative;display:flex;justify-content:center;align-self:center;width:100%;max-width:100%;flex-direction:column;height:320px"><iframe class="notion-asset-object-fit" src="https://platform.twitter.com/embed/Tweet.html?id=2081940052154933696" title="iframe embed" frameBorder="0" allowfullscreen="" loading="lazy" scrolling="auto"></iframe></div></figure><figure class="notion-asset-wrapper notion-asset-wrapper-embed notion-block-b8c453aee0d54291979bc0c97f12f0a1"><div style="position:relative;display:flex;justify-content:center;align-self:center;width:100%;max-width:100%;flex-direction:column;height:320px"><iframe class="notion-asset-object-fit" src="https://platform.twitter.com/embed/Tweet.html?id=2081899343091843463" title="iframe embed" frameBorder="0" allowfullscreen="" loading="lazy" scrolling="auto"></iframe></div></figure><h4 class="notion-h notion-h3 notion-h-indent-1 notion-block-8f62335aef1a4e4fa8b9933199c94399" data-id="8f62335aef1a4e4fa8b9933199c94399"><span><div id="8f62335aef1a4e4fa8b9933199c94399" class="notion-header-anchor"></div><a class="notion-hash-link" href="#8f62335aef1a4e4fa8b9933199c94399" title="3. Peter Yang @petergyang"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">3. Peter Yang @petergyang</span></span></h4><div class="notion-text notion-block-6540c0f365c14288a77dbabbebac34bf">Peter Yang 转述 OpenAI DevEx Jason 的 Codex 使用案例：人还在骑车时，通过手机远程让 Codex 用 computer use 修改 launch video、导出并发到 Slack；之后又让 Codex 每 30 分钟检查 Slack 反馈并生成 V2 / V3 / V4，直到视频过审。这个例子很有信息量，因为 Codex 承担的是跨工具、跨反馈循环的交付任务，不只是写代码。他另外两条补充了完整访谈和文字版 prompts。</div><figure class="notion-asset-wrapper notion-asset-wrapper-embed notion-block-12928a5de3e4433d941647ade2091625"><div style="position:relative;display:flex;justify-content:center;align-self:center;width:100%;max-width:100%;flex-direction:column;height:320px"><iframe class="notion-asset-object-fit" src="https://platform.twitter.com/embed/Tweet.html?id=2081775399097549083" title="iframe embed" frameBorder="0" allowfullscreen="" loading="lazy" scrolling="auto"></iframe></div></figure><figure class="notion-asset-wrapper notion-asset-wrapper-embed notion-block-a73b12146b3241d8948bf5f8180c357a"><div style="position:relative;display:flex;justify-content:center;align-self:center;width:100%;max-width:100%;flex-direction:column;height:320px"><iframe class="notion-asset-object-fit" src="https://platform.twitter.com/embed/Tweet.html?id=2081767570198401263" title="iframe embed" frameBorder="0" allowfullscreen="" loading="lazy" scrolling="auto"></iframe></div></figure><figure class="notion-asset-wrapper notion-asset-wrapper-embed notion-block-84f8c66bd21c459c9750f3a2fac35939"><div style="position:relative;display:flex;justify-content:center;align-self:center;width:100%;max-width:100%;flex-direction:column;height:320px"><iframe class="notion-asset-object-fit" src="https://platform.twitter.com/embed/Tweet.html?id=2081767558408175867" title="iframe embed" frameBorder="0" allowfullscreen="" loading="lazy" scrolling="auto"></iframe></div></figure><h4 class="notion-h notion-h3 notion-h-indent-1 notion-block-13f8e72fab4e43f288febfa4c25a7c87" data-id="13f8e72fab4e43f288febfa4c25a7c87"><span><div id="13f8e72fab4e43f288febfa4c25a7c87" class="notion-header-anchor"></div><a class="notion-hash-link" href="#13f8e72fab4e43f288febfa4c25a7c87" title="4. Nan Yu @thenanyu"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">4. Nan Yu @thenanyu</span></span></h4><div class="notion-text notion-block-a83b1e36fbd84f45a4844e199cddf1f4">Nan Yu 一条是轻量动态，另一条更像产品哲学：如果团队有聪明人，就应该让他们把自己的产品打磨得更好、更好用，同时愿意为好工具付费。放在 Linear 的背景下看，这不是 AI 新闻，但和今天 “AI-native 工作界面” 主线有关：当大家都在追 agent，产品基础体验和品味仍然不能被跳过。</div><figure class="notion-asset-wrapper notion-asset-wrapper-embed notion-block-e942ca6d99dc45399353baead9b302db"><div style="position:relative;display:flex;justify-content:center;align-self:center;width:100%;max-width:100%;flex-direction:column;height:320px"><iframe class="notion-asset-object-fit" src="https://platform.twitter.com/embed/Tweet.html?id=2081926688250691884" title="iframe embed" frameBorder="0" allowfullscreen="" loading="lazy" scrolling="auto"></iframe></div></figure><figure class="notion-asset-wrapper notion-asset-wrapper-embed notion-block-75de1ddce61c490a9b395c436517ce07"><div style="position:relative;display:flex;justify-content:center;align-self:center;width:100%;max-width:100%;flex-direction:column;height:320px"><iframe class="notion-asset-object-fit" src="https://platform.twitter.com/embed/Tweet.html?id=2081768780045156358" title="iframe embed" frameBorder="0" allowfullscreen="" loading="lazy" scrolling="auto"></iframe></div></figure><h4 class="notion-h notion-h3 notion-h-indent-1 notion-block-e5ed576476e64c3aa2ac98ee2859491b" data-id="e5ed576476e64c3aa2ac98ee2859491b"><span><div id="e5ed576476e64c3aa2ac98ee2859491b" class="notion-header-anchor"></div><a class="notion-hash-link" href="#e5ed576476e64c3aa2ac98ee2859491b" title="5. Madhu Guru @realmadhuguru"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">5. Madhu Guru @realmadhuguru</span></span></h4><div class="notion-text notion-block-a3efffdcc7ad45478cafab39948906a0">Madhu Guru 讲的是 product review 的质量。他认为好的 product review 是把几个月的学习压缩到一小时，让房间里的人模拟市场对想法的反应；坏的 review 会退化成状态更新、领导曝光和跨部门对齐。这条不是模型动态，但对 AI 产品团队很实用，因为 agent 工具越强，团队越需要高质量判断来决定要做什么。</div><figure class="notion-asset-wrapper notion-asset-wrapper-embed notion-block-50fa22611b0e44209d5701e43ae9a9e2"><div style="position:relative;display:flex;justify-content:center;align-self:center;width:100%;max-width:100%;flex-direction:column;height:320px"><iframe class="notion-asset-object-fit" src="https://platform.twitter.com/embed/Tweet.html?id=2081781952437486052" title="iframe embed" frameBorder="0" allowfullscreen="" loading="lazy" scrolling="auto"></iframe></div></figure><h4 class="notion-h notion-h3 notion-h-indent-1 notion-block-790cdcfc712947649b2b906c624bff1f" data-id="790cdcfc712947649b2b906c624bff1f"><span><div id="790cdcfc712947649b2b906c624bff1f" class="notion-header-anchor"></div><a class="notion-hash-link" href="#790cdcfc712947649b2b906c624bff1f" title="6. Amjad Masad @amasad"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">6. Amjad Masad @amasad</span></span></h4><div class="notion-text notion-block-2ad9b35d556146688353af22e4254bb4">Amjad Masad 把 AI agent 看成新一代“探索工具”。他的类比是：祖先探索地球，后来探索太空，而这一代人可能探索 computational universe，也就是算法、程序、证明和设计的巨大空间。这个判断和 Replit 的方向一致：agent 不只是自动化已有工作，还会扩展人类能搜索和尝试的设计空间。</div><figure class="notion-asset-wrapper notion-asset-wrapper-embed notion-block-c856e6e727c0461990a5b44b02134514"><div style="position:relative;display:flex;justify-content:center;align-self:center;width:100%;max-width:100%;flex-direction:column;height:320px"><iframe class="notion-asset-object-fit" src="https://platform.twitter.com/embed/Tweet.html?id=2082000490066592127" title="iframe embed" frameBorder="0" allowfullscreen="" loading="lazy" scrolling="auto"></iframe></div></figure><h4 class="notion-h notion-h3 notion-h-indent-1 notion-block-8bcfb609131f4011848fff4541a0fcc8" data-id="8bcfb609131f4011848fff4541a0fcc8"><span><div id="8bcfb609131f4011848fff4541a0fcc8" class="notion-header-anchor"></div><a class="notion-hash-link" href="#8bcfb609131f4011848fff4541a0fcc8" title="7. Guillermo Rauch @rauchg"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">7. Guillermo Rauch @rauchg</span></span></h4><div class="notion-text notion-block-aa074723a882482a872db14947bc8719">Guillermo Rauch 今天两条都很关键：一条用 Vercel benchmark 比较 cybersecurity AI model 的 price-performance，认为 Grok 4.5 在性价比上很强，Sol 仍是 frontier；另一条从 Kimi paper 里提炼 agent 安全边界，认为 container-level isolation 不够，Firecracker microVM 这类隔离更适合 agent 运行。他还有一条 EU/acc 轻量转发，也按全量要求保留。这里最重要的是：模型能力、成本和运行边界正在一起被比较。</div><figure class="notion-asset-wrapper notion-asset-wrapper-embed notion-block-f6427f1c537548df857ed5330830efea"><div style="position:relative;display:flex;justify-content:center;align-self:center;width:100%;max-width:100%;flex-direction:column;height:320px"><iframe class="notion-asset-object-fit" src="https://platform.twitter.com/embed/Tweet.html?id=2081852481517318560" title="iframe embed" frameBorder="0" allowfullscreen="" loading="lazy" scrolling="auto"></iframe></div></figure><figure class="notion-asset-wrapper notion-asset-wrapper-embed notion-block-281842dbc8a44483928a5b35707f3800"><div style="position:relative;display:flex;justify-content:center;align-self:center;width:100%;max-width:100%;flex-direction:column;height:320px"><iframe class="notion-asset-object-fit" src="https://platform.twitter.com/embed/Tweet.html?id=2081845695112446364" title="iframe embed" frameBorder="0" allowfullscreen="" loading="lazy" scrolling="auto"></iframe></div></figure><figure class="notion-asset-wrapper notion-asset-wrapper-embed notion-block-a05e7ff3675d4c0790dce377c14269c6"><div style="position:relative;display:flex;justify-content:center;align-self:center;width:100%;max-width:100%;flex-direction:column;height:320px"><iframe class="notion-asset-object-fit" src="https://platform.twitter.com/embed/Tweet.html?id=2081842439304995169" title="iframe embed" frameBorder="0" allowfullscreen="" loading="lazy" scrolling="auto"></iframe></div></figure><h4 class="notion-h notion-h3 notion-h-indent-1 notion-block-b29d3742128349a4a73ca3fd80a97d1b" data-id="b29d3742128349a4a73ca3fd80a97d1b"><span><div id="b29d3742128349a4a73ca3fd80a97d1b" class="notion-header-anchor"></div><a class="notion-hash-link" href="#b29d3742128349a4a73ca3fd80a97d1b" title="8. Aaron Levie @levie"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">8. Aaron Levie @levie</span></span></h4><div class="notion-text notion-block-c6c0049bad7441c4820922197f1bf8c6">Aaron Levie 继续反对“AI 必然导致大规模岗位消失”的简单叙事。他看到的企业情况是仍在招聘，只是岗位方向发生变化：更多工程师去做过去做不了的问题，更多销售去深化客户关系，也更多 internal FDE 帮公司部署 AI。他的判断是，只把 AI 用来降本的公司会输给用 AI 更好服务客户、推动突破的公司。另一条关于 K3 weights，也和开放模型动态有关。</div><figure class="notion-asset-wrapper notion-asset-wrapper-embed notion-block-c0bd36338863482890748e953cb23a98"><div style="position:relative;display:flex;justify-content:center;align-self:center;width:100%;max-width:100%;flex-direction:column;height:320px"><iframe class="notion-asset-object-fit" src="https://platform.twitter.com/embed/Tweet.html?id=2081930301752942703" title="iframe embed" frameBorder="0" allowfullscreen="" loading="lazy" scrolling="auto"></iframe></div></figure><figure class="notion-asset-wrapper notion-asset-wrapper-embed notion-block-4ddd0b9020634ce69ca15a5b50ff5744"><div style="position:relative;display:flex;justify-content:center;align-self:center;width:100%;max-width:100%;flex-direction:column;height:320px"><iframe class="notion-asset-object-fit" src="https://platform.twitter.com/embed/Tweet.html?id=2081760710108012702" title="iframe embed" frameBorder="0" allowfullscreen="" loading="lazy" scrolling="auto"></iframe></div></figure><h4 class="notion-h notion-h3 notion-h-indent-1 notion-block-c6a3159969c44db881626bbc3c991b6f" data-id="c6a3159969c44db881626bbc3c991b6f"><span><div id="c6a3159969c44db881626bbc3c991b6f" class="notion-header-anchor"></div><a class="notion-hash-link" href="#c6a3159969c44db881626bbc3c991b6f" title="9. Matt Turck @mattturck"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">9. Matt Turck @mattturck</span></span></h4><div class="notion-text notion-block-96ecca6c341845d0a6a9a161a7df7d2e">Matt Turck 今天这条是 VC 圈自嘲，和 AI 主线关系不强，但属于原始抓取内容，保留。它大意是研究显示不到 40% 的 VC 有成功投资，而每个 VC 都觉得自己在那 40% 里。</div><figure class="notion-asset-wrapper notion-asset-wrapper-embed notion-block-8b574dabebb7461cb3f8bdd07668c458"><div style="position:relative;display:flex;justify-content:center;align-self:center;width:100%;max-width:100%;flex-direction:column;height:320px"><iframe class="notion-asset-object-fit" src="https://platform.twitter.com/embed/Tweet.html?id=2081679801769668980" title="iframe embed" frameBorder="0" allowfullscreen="" loading="lazy" scrolling="auto"></iframe></div></figure><h4 class="notion-h notion-h3 notion-h-indent-1 notion-block-98e20cd6a4694c74bc3bdaf4a5c9d8fd" data-id="98e20cd6a4694c74bc3bdaf4a5c9d8fd"><span><div id="98e20cd6a4694c74bc3bdaf4a5c9d8fd" class="notion-header-anchor"></div><a class="notion-hash-link" href="#98e20cd6a4694c74bc3bdaf4a5c9d8fd" title="10. Zara Zhang @zarazhangrui"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">10. Zara Zhang @zarazhangrui</span></span></h4><div class="notion-text notion-block-1b079e3413a4479b8046a1ff0effcaed">Zara Zhang 今天两条偏内容创作和个人工作状态。一条分享她如何不断获得内容想法，另一条是“你想要的 magic 在你正在回避的工作里”。它们不是模型或产品发布，但和 builder 表达有关：AI 工具再强，持续输出和面对难任务仍然是底层能力。</div><figure class="notion-asset-wrapper notion-asset-wrapper-embed notion-block-34f05df05aa648fd9ae656af33a9883f"><div style="position:relative;display:flex;justify-content:center;align-self:center;width:100%;max-width:100%;flex-direction:column;height:320px"><iframe class="notion-asset-object-fit" src="https://platform.twitter.com/embed/Tweet.html?id=2081983750658044079" title="iframe embed" frameBorder="0" allowfullscreen="" loading="lazy" scrolling="auto"></iframe></div></figure><figure class="notion-asset-wrapper notion-asset-wrapper-embed notion-block-2f03ac157acc42e583f72e469c33fabc"><div style="position:relative;display:flex;justify-content:center;align-self:center;width:100%;max-width:100%;flex-direction:column;height:320px"><iframe class="notion-asset-object-fit" src="https://platform.twitter.com/embed/Tweet.html?id=2081976736854737164" title="iframe embed" frameBorder="0" allowfullscreen="" loading="lazy" scrolling="auto"></iframe></div></figure><h4 class="notion-h notion-h3 notion-h-indent-1 notion-block-c49966a1aee7412390daf786e881cc53" data-id="c49966a1aee7412390daf786e881cc53"><span><div id="c49966a1aee7412390daf786e881cc53" class="notion-header-anchor"></div><a class="notion-hash-link" href="#c49966a1aee7412390daf786e881cc53" title="11. Nikunj Kothari @nikunj"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">11. Nikunj Kothari @nikunj</span></span></h4><div class="notion-text notion-block-fb81d61fad7847658f18fabd97dac654">Nikunj Kothari 今天最相关的是第一条：他把 Claude Code 当作两周旅行的主要界面，然后让它做完整复盘，看看下次怎么改进。这是一个很好的“生活工作流 agent”例子。其他两条一条是 Micro 电动车的发现，一条是社交玩笑，和 AI 主线关系不强，但按全量要求保留。</div><figure class="notion-asset-wrapper notion-asset-wrapper-embed notion-block-23e37174165342b9adaf4649cd206c38"><div style="position:relative;display:flex;justify-content:center;align-self:center;width:100%;max-width:100%;flex-direction:column;height:320px"><iframe class="notion-asset-object-fit" src="https://platform.twitter.com/embed/Tweet.html?id=2081992618649547100" title="iframe embed" frameBorder="0" allowfullscreen="" loading="lazy" scrolling="auto"></iframe></div></figure><figure class="notion-asset-wrapper notion-asset-wrapper-embed notion-block-00a8802128cc47f283fc9c3954c798d9"><div style="position:relative;display:flex;justify-content:center;align-self:center;width:100%;max-width:100%;flex-direction:column;height:320px"><iframe class="notion-asset-object-fit" src="https://platform.twitter.com/embed/Tweet.html?id=2081805464757485706" title="iframe embed" frameBorder="0" allowfullscreen="" loading="lazy" scrolling="auto"></iframe></div></figure><figure class="notion-asset-wrapper notion-asset-wrapper-embed notion-block-a1ca860e43614a7da5359e5cdea58c82"><div style="position:relative;display:flex;justify-content:center;align-self:center;width:100%;max-width:100%;flex-direction:column;height:320px"><iframe class="notion-asset-object-fit" src="https://platform.twitter.com/embed/Tweet.html?id=2081750712761852341" title="iframe embed" frameBorder="0" allowfullscreen="" loading="lazy" scrolling="auto"></iframe></div></figure><h4 class="notion-h notion-h3 notion-h-indent-1 notion-block-20dde147cb2747688daaab3792957e49" data-id="20dde147cb2747688daaab3792957e49"><span><div id="20dde147cb2747688daaab3792957e49" class="notion-header-anchor"></div><a class="notion-hash-link" href="#20dde147cb2747688daaab3792957e49" title="12. Peter Steinberger @steipete"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">12. Peter Steinberger @steipete</span></span></h4><div class="notion-text notion-block-cb1d460ca1f54784acbfe1ec5b7b5333">Peter Steinberger 今天有一条非常 agent-native：他的 agent 报了一个 bug，对方的 agent 当晚修掉，说明 agent-to-agent 协作已经开始出现在开发者日常里。另一条是关于安全工作的外界信任问题，和今天 agent 安全边界主线相呼应；还有一条活动见面信息，按全量记录保留。</div><figure class="notion-asset-wrapper notion-asset-wrapper-embed notion-block-919aae24ac874719ad2118dd8eca6fde"><div style="position:relative;display:flex;justify-content:center;align-self:center;width:100%;max-width:100%;flex-direction:column;height:320px"><iframe class="notion-asset-object-fit" src="https://platform.twitter.com/embed/Tweet.html?id=2081865727443902654" title="iframe embed" frameBorder="0" allowfullscreen="" loading="lazy" scrolling="auto"></iframe></div></figure><figure class="notion-asset-wrapper notion-asset-wrapper-embed notion-block-a9e5694196114f8cb351d1345560cc5f"><div style="position:relative;display:flex;justify-content:center;align-self:center;width:100%;max-width:100%;flex-direction:column;height:320px"><iframe class="notion-asset-object-fit" src="https://platform.twitter.com/embed/Tweet.html?id=2081790109415002468" title="iframe embed" frameBorder="0" allowfullscreen="" loading="lazy" scrolling="auto"></iframe></div></figure><figure class="notion-asset-wrapper notion-asset-wrapper-embed notion-block-1fe06d741a684c5da10be3a35634198f"><div style="position:relative;display:flex;justify-content:center;align-self:center;width:100%;max-width:100%;flex-direction:column;height:320px"><iframe class="notion-asset-object-fit" src="https://platform.twitter.com/embed/Tweet.html?id=2081767828278170002" title="iframe embed" frameBorder="0" allowfullscreen="" loading="lazy" scrolling="auto"></iframe></div></figure><h4 class="notion-h notion-h3 notion-h-indent-1 notion-block-0fda76ae9f1b4cca9941821edde61bcc" data-id="0fda76ae9f1b4cca9941821edde61bcc"><span><div id="0fda76ae9f1b4cca9941821edde61bcc" class="notion-header-anchor"></div><a class="notion-hash-link" href="#0fda76ae9f1b4cca9941821edde61bcc" title="13. Dan Shipper @danshipper"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">13. Dan Shipper @danshipper</span></span></h4><div class="notion-text notion-block-ac9cd4c6128543c4bfbfbb9dda03cf52">Dan Shipper 这条是关于 rare books 的非 AI 内容，信息量不大，但属于本次 follow-builders 原始抓取，保留。</div><figure class="notion-asset-wrapper notion-asset-wrapper-embed notion-block-d333ab0aa687403e816c648fb9121534"><div style="position:relative;display:flex;justify-content:center;align-self:center;width:100%;max-width:100%;flex-direction:column;height:320px"><iframe class="notion-asset-object-fit" src="https://platform.twitter.com/embed/Tweet.html?id=2081754482568835152" title="iframe embed" frameBorder="0" allowfullscreen="" loading="lazy" scrolling="auto"></iframe></div></figure><h4 class="notion-h notion-h3 notion-h-indent-1 notion-block-500b34c88e8c407c8583c5dc80babd87" data-id="500b34c88e8c407c8583c5dc80babd87"><span><div id="500b34c88e8c407c8583c5dc80babd87" class="notion-header-anchor"></div><a class="notion-hash-link" href="#500b34c88e8c407c8583c5dc80babd87" title="14. Sam Altman @sama"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">14. Sam Altman @sama</span></span></h4><div class="notion-text notion-block-e2f69bd472904553b7d79171b3581cd0">Sam Altman 今天只有一个很短的“wrong”回应，JSON 里没有更多上下文。按规则不补外部信息，只保留原帖。</div><figure class="notion-asset-wrapper notion-asset-wrapper-embed notion-block-69605df78c5f447fb5483b49fd291b91"><div style="position:relative;display:flex;justify-content:center;align-self:center;width:100%;max-width:100%;flex-direction:column;height:320px"><iframe class="notion-asset-object-fit" src="https://platform.twitter.com/embed/Tweet.html?id=2081832600591892712" title="iframe embed" frameBorder="0" allowfullscreen="" loading="lazy" scrolling="auto"></iframe></div></figure><h3 class="notion-h notion-h2 notion-h-indent-0 notion-block-45d57b103e764c90abf9f1157d84c226" data-id="45d57b103e764c90abf9f1157d84c226"><span><div id="45d57b103e764c90abf9f1157d84c226" class="notion-header-anchor"></div><a class="notion-hash-link" href="#45d57b103e764c90abf9f1157d84c226" title="Blog 全量记录"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">Blog 全量记录</span></span></h3><div class="notion-text notion-block-8e8d5322b12c45d6ab8ea5fa4048b926">今天 JSON 中没有新增 blog。</div><h3 class="notion-h notion-h2 notion-h-indent-0 notion-block-2550d6329ed84864a4ace956119d3c75" data-id="2550d6329ed84864a4ace956119d3c75"><span><div id="2550d6329ed84864a4ace956119d3c75" class="notion-header-anchor"></div><a class="notion-hash-link" href="#2550d6329ed84864a4ace956119d3c75" title="Podcast 全量记录"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">Podcast 全量记录</span></span></h3><h4 class="notion-h notion-h3 notion-h-indent-1 notion-block-bb27be96de3b471da277b1e520def3f2" data-id="bb27be96de3b471da277b1e520def3f2"><span><div id="bb27be96de3b471da277b1e520def3f2" class="notion-header-anchor"></div><a class="notion-hash-link" href="#bb27be96de3b471da277b1e520def3f2" title="1. AI &amp; I by Every：The Founder of a $1.5B AI Company on What Comes After the First Wave of AI Apps"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">1. AI &amp; I by Every：The Founder of a $1.5B AI Company on What Comes After the First Wave of AI Apps</span></span></h4><div class="notion-text notion-block-b0495c7e9de94bb8a8ce82e3f1e16e5c">Takeaway：meeting notes 只是第一波 AI app 的入口，不是终点。Granola CEO 的核心判断是，真正的大机会在于重新发明人和 AI 一起工作的界面。他认为 AI-native 工作会从单点功能走向更大的工作系统：一边是 Slack 里公开、异步、多人协作的 agent delegation，另一边是 Codex、Claude desktop 这类“同屏工作”的深度协作空间。访谈里还有一个很实在的创业视角：即使产品已经跑起来，创业仍然像 “knife fights”；公司从 12 人长到 60-70 人后，如何保持产品的 soul、taste 和一致性，反而更难。对 AI 产品来说，复制 meeting notes 按钮不一定会杀死原产品，因为真正的竞争还没到终局，下一代工作界面才是关键。来源：<a target="_blank" rel="noopener noreferrer" class="notion-link" href="https://www.youtube.com/playlist?list=PLuMcoKK9mKgHtW_o9h5sGO2vXrffKHwJL">YouTube playlist 原始链接</a></div><h3 class="notion-h notion-h2 notion-h-indent-0 notion-block-40dfdeb6d6004585997249bc26736a04" data-id="40dfdeb6d6004585997249bc26736a04"><span><div id="40dfdeb6d6004585997249bc26736a04" class="notion-header-anchor"></div><a class="notion-hash-link" href="#40dfdeb6d6004585997249bc26736a04" title="今日沉淀结论"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">今日沉淀结论</span></span></h3><div class="notion-text notion-block-b4b724949d334c5b8ca10d227ab6db49">今天所有内容放在一起看，agent 正在离开 demo 阶段，进入真实工作流、产品组织、成本评估和安全边界。真正值得跟的是三个信号：任务成本而不是 token 成本；微隔离、安全审计和运行边界；以及 AI-native 工作界面从单人聊天变成多人协作系统。</div><div class="notion-text notion-block-60fe85c29ccc49b29f300b8c3788320f">Generated through the Follow Builders skill: <a target="_blank" rel="noopener noreferrer" class="notion-link" href="https://github.com/zarazhangrui/follow-builders">https://github.com/zarazhangrui/follow-builders</a></div></main></div>]]></content:encoded>
        </item>
        <item>
            <title><![CDATA[Skill Self-Play 深度分析 - 2026-07-28]]></title>
            <link>http://easyai.fyi/article/skill-self-play-deep-analysis-2026-07-28</link>
            <guid>http://easyai.fyi/article/skill-self-play-deep-analysis-2026-07-28</guid>
            <pubDate>Tue, 28 Jul 2026 00:00:00 GMT</pubDate>
            <description><![CDATA[Skill Self-Play 用可进化 skill 生成、验证并调度训练任务，让 Agent 的课程跟着能力边界变化。]]></description>
            <content:encoded><![CDATA[<div id="notion-article" class="mx-auto overflow-hidden "><main class="notion light-mode notion-page notion-block-3ab5eac3813c815aa1badea0728ccc65"><div class="notion-viewport"></div><div class="notion-collection-page-properties"></div><blockquote class="notion-quote notion-block-9903c7142d944465bd00f72d37ea260c"><div><b>一句话判断</b>：Skill Self-Play 把 skill 从“运行时给 Agent 看的说明书”搬进训练闭环，让它同时负责出题、验题和记录课程进度。这个位置变化很有意思，但论文目前只在工具调用与逻辑题上成立，离通用 Agent 自我进化还有一段距离。</div></blockquote><h3 class="notion-h notion-h2 notion-h-indent-0 notion-block-5d411e3ad7b5480c8d2eff149f456fcc" data-id="5d411e3ad7b5480c8d2eff149f456fcc"><span><div id="5d411e3ad7b5480c8d2eff149f456fcc" class="notion-header-anchor"></div><a class="notion-hash-link" href="#5d411e3ad7b5480c8d2eff149f456fcc" title="原始材料"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">原始材料</span></span></h3><ul class="notion-list notion-list-disc notion-block-0dcac25f40d24b6e9da7f22983cb792c"><li><a target="_blank" rel="noopener noreferrer" class="notion-link" href="https://arxiv.org/abs/2607.22529">论文页面：arXiv:2607.22529</a></li></ul><ul class="notion-list notion-list-disc notion-block-99df7e84238e4b60aab9b317a36ca631"><li><a target="_blank" rel="noopener noreferrer" class="notion-link" href="https://arxiv.org/pdf/2607.22529">论文 PDF</a></li></ul><ul class="notion-list notion-list-disc notion-block-5fb7f4494e114e1da43598992b22c78d"><li><a target="_blank" rel="noopener noreferrer" class="notion-link" href="https://github.com/Qwen-Applications/skill-self-play">开源仓库：Qwen-Applications/skill-self-play</a></li></ul><ul class="notion-list notion-list-disc notion-block-a423ad76fb6046e5ad509bdfb6a25977"><li><a target="_blank" rel="noopener noreferrer" class="notion-link" href="https://huggingface.co/papers/2607.22529">Hugging Face 论文页</a></li></ul><div class="notion-text notion-block-fd3fe8683a5d44d8b1f198c31d090f93">论文由阿里巴巴 Qwen 大模型应用团队等机构的研究者完成，2026 年 7 月 24 日提交。代码采用 Apache 2.0 许可证，于 7 月 27 日公开。</div><h3 class="notion-h notion-h2 notion-h-indent-0 notion-block-fe3234c8549e4395ad430d2c301d7ff8" data-id="fe3234c8549e4395ad430d2c301d7ff8"><span><div id="fe3234c8549e4395ad430d2c301d7ff8" class="notion-header-anchor"></div><a class="notion-hash-link" href="#fe3234c8549e4395ad430d2c301d7ff8" title="为什么今天选它"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">为什么今天选它</span></span></h3><div class="notion-text notion-block-f04689bfd2414d25a714f8a5da994939">7 月 27 日的 Hugging Face Daily Papers 里，Skill Self-Play 排到<a target="_blank" rel="noopener noreferrer" class="notion-link" href="https://huggingface.co/papers/2607.22529">当日第 2</a>。同期几个值得看的对象各有侧重：Molt 主要解决 Agent 强化学习框架太重的问题；Multi-Head Latent Control 研究模型何时调用工具、澄清、拒答或升级；IDEAgent 研究科研创意的质量与多样性搜索。Skill Self-Play 更适合今天拆，因为它把 skills、self-play、验证器、课程学习和训练闭环放进了同一套系统，而且代码已经公开。</div><div class="notion-text notion-block-570c38a5a80b42d386f9d19da78dd8cf">【智汇AI】7 月 14 日写过 SkillOpt。两者看起来都在“让 skill 进化”，实际位置不同：SkillOpt 把 skill 当作运行时可加载的外部能力；Skill Self-Play 把 skill 当作训练时的课程接口。后者训练结束后，solver 推理时不再加载这些 skill。这个区别决定了它更像训练方法，而不是一个可直接接入产品的 Agent 插件系统。</div><h3 class="notion-h notion-h2 notion-h-indent-0 notion-block-fa6810a336b94c29959dfb4fbc311009" data-id="fa6810a336b94c29959dfb4fbc311009"><span><div id="fa6810a336b94c29959dfb4fbc311009" class="notion-header-anchor"></div><a class="notion-hash-link" href="#fa6810a336b94c29959dfb4fbc311009" title="它在解决什么问题"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">它在解决什么问题</span></span></h3><div class="notion-text notion-block-c2ec7afa469a4401ab7e2d27f8b240a6">自我进化训练常卡在两头。</div><div class="notion-text notion-block-1108c59674294a589891395006692d24">一头是固定环境。代码执行器、游戏模拟器和结构化 API 能给出可靠答案，但任务边界早已写死。模型变强后，环境未必能继续提供合适的新题。</div><div class="notion-text notion-block-fec17b722e984e21a5d7e1de34871deb">另一头是开放生成。让模型自己出题可以扩大覆盖面，可一旦题目本身有歧义、答案错误或无法验收，错误奖励就会进入下一轮训练。普通格式检查只能拦住明显坏样本，拦不住“格式正确、逻辑有问题”的题。</div><div class="notion-text notion-block-da6934072ab74688a10517c46fe88e21">Skill-SP 的做法是把 skill 定义成一份可执行的任务模式。它不只写“怎么做”，还带着出题约束、示例、验证器和使用统计。系统用这些结构先约束题目怎么生成，再检查题目是否有效；与此同时，它保留一条不受现有 skill 约束的探索通道，用来发现新模式。</div><figure class="notion-asset-wrapper notion-asset-wrapper-image notion-block-161d283924eb4500b652c4d0e2492555"><div style="position:relative;display:flex;justify-content:center;align-self:center;width:100%;max-width:100%;flex-direction:column"><img style="object-fit:cover" src="https://raw.githubusercontent.com/Qwen-Applications/skill-self-play/main/assets/skill_sp_framework.png?t=161d2839-24eb-4500-b652-c4d0e2492555" alt="notion image" loading="lazy" decoding="async"/></div></figure><div class="notion-text notion-block-e299bea09f9c4e5290b9a151eea6ea33"><em>图 1：论文 Figure 3，也是开源仓库的官方流程图。skill 库先指导 proposer 出题，验证通过后按难度组装课程；失败记录、新样本和成功率再回流给 controller。来源：</em><em><a target="_blank" rel="noopener noreferrer" class="notion-link" href="https://github.com/Qwen-Applications/skill-self-play/blob/main/assets/skill_sp_framework.png">官方仓库</a></em><em>。</em></div><h3 class="notion-h notion-h2 notion-h-indent-0 notion-block-662b7601ca2943cbbf724f88a90c7fe9" data-id="662b7601ca2943cbbf724f88a90c7fe9"><span><div id="662b7601ca2943cbbf724f88a90c7fe9" class="notion-header-anchor"></div><a class="notion-hash-link" href="#662b7601ca2943cbbf724f88a90c7fe9" title="系统怎么运作"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">系统怎么运作</span></span></h3><h4 class="notion-h notion-h3 notion-h-indent-1 notion-block-78557bd10f944213adb2184d43415b62" data-id="78557bd10f944213adb2184d43415b62"><span><div id="78557bd10f944213adb2184d43415b62" class="notion-header-anchor"></div><a class="notion-hash-link" href="#78557bd10f944213adb2184d43415b62" title="1. skill 是训练资产，不是提示词片段"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">1. skill 是训练资产，不是提示词片段</span></span></h4><div class="notion-text notion-block-2cde7c11700b4dd1b06e5b79d2ff2a1c">公开代码里的初始工具调用库有 15 个 skill package。每个目录包含 <code class="notion-inline-code">SKILL.md</code>、规则、生成提示、示例、solver 侧规则、统计文件和可执行 validator。比如“地理编码结果转坐标服务”这一项会要求下一次调用复用已有经纬度，并禁止重复调用 geocoding；validator 会检查工具名称、必填参数和参数是否来自已有观察。<a target="_blank" rel="noopener noreferrer" class="notion-link" href="https://github.com/Qwen-Applications/skill-self-play/tree/main/tool_call/packages">仓库中的 package 说明</a>能看到这套结构。</div><div class="notion-text notion-block-8d27516c06304f2fb136aaf7480438e8">这说明论文里的 skill 不是一段泛化建议。它更接近“任务模板 + 局部验收器 + 课程状态”。加载也分层：发现时先读元数据，选中后再读规则，只有需要完整诊断时才展开示例。这套 progressive disclosure 与常见 Agent Skill 的工程形式相似，只是服务对象变成了出题模型。</div><h4 class="notion-h notion-h3 notion-h-indent-1 notion-block-dc854863a1d5487cb002c6fa331d1bd4" data-id="dc854863a1d5487cb002c6fa331d1bd4"><span><div id="dc854863a1d5487cb002c6fa331d1bd4" class="notion-header-anchor"></div><a class="notion-hash-link" href="#dc854863a1d5487cb002c6fa331d1bd4" title="2. proposer、solver、controller 各管一件事"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">2. proposer、solver、controller 各管一件事</span></span></h4><div class="notion-text notion-block-52a2f6a5f12a445f9e57b79a368f1b57">Proposer 负责生成题目和隐藏的验收合同。Solver 解题，也被当成当前能力的经验评估器。Controller 不直接解题，它根据生成失败、题目新颖度、solver 成功率和任务效用来维护 skill 库。</div><div class="notion-text notion-block-adc3a48900c044eaad91b637ed447e4d">每轮同时走两条数据通道：</div><ul class="notion-list notion-list-disc notion-block-e0c3b3820d0d4455afa7aa84c2795d6b"><li>skill stream 从库里抽取一项 skill，按它的结构生成题目；</li></ul><ul class="notion-list notion-list-disc notion-block-b81e1c4ed5af4beaa2eca18a48d595c2"><li>exploration stream 不带 skill，自由探索新题型。</li></ul><div class="notion-text notion-block-bae4107378624dbb8e0cbedc0473a7ed">只用 skill 会过度收窄，只做自由探索又容易产生噪声。论文把两路有效样本混在一起，默认各占一半。</div><h4 class="notion-h notion-h3 notion-h-indent-1 notion-block-7687f05e11084c0dae03e3693ff04f26" data-id="7687f05e11084c0dae03e3693ff04f26"><span><div id="7687f05e11084c0dae03e3693ff04f26" class="notion-header-anchor"></div><a class="notion-hash-link" href="#7687f05e11084c0dae03e3693ff04f26" title="3. 先判断“题是不是题”，再判断“难不难”"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">3. 先判断“题是不是题”，再判断“难不难”</span></span></h4><div class="notion-text notion-block-335a44e441374bffb90f5d3697761f5b">每个任务被表示成用户可见问题和机器可读的隐藏合同。工具调用任务检查函数名、参数与格式；逻辑题使用确定性求解器确认约束成立且答案唯一。对 skill stream，系统还会运行 package 自带的 validator，并用 solver 的多次探测结果检查参考答案是否一致。</div><div class="notion-text notion-block-9431d33eb21144bfa30db94639d3a7ca">通过有效性门槛后，系统才计算 frontier reward。它偏好 solver 成功率接近 0.5 的题：太容易没有训练价值，太难又拿不到学习信号。这个门槛很重要，因为单纯奖励“把 solver 难住”会鼓励 proposer 造坏题。</div><h4 class="notion-h notion-h3 notion-h-indent-1 notion-block-dee52b41143a4fb887e36949c74dc27e" data-id="dee52b41143a4fb887e36949c74dc27e"><span><div id="dee52b41143a4fb887e36949c74dc27e" class="notion-header-anchor"></div><a class="notion-hash-link" href="#dee52b41143a4fb887e36949c74dc27e" title="4. 题库和模型一起变"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">4. 题库和模型一起变</span></span></h4><div class="notion-text notion-block-162ac21248e64b87bdbc75406f794ce7">Proposer 与 solver 都用 GRPO 更新。Solver 变强后，原来合适的题会逐渐变简单；controller 会降低这类 skill 的抽样权重，必要时归档。探索通道里出现有效且有难度的新模式时，controller 把它归纳成新 package。论文实现还会检查 package 是否完整，并用词面相似度过滤重复 skill。</div><div class="notion-text notion-block-ed4b3898b8f24164a9aaae11f83da1cd">训练共跑 5 轮。工具调用每轮选 8,000 道题，逻辑推理每轮选 1,920 道题。所有实验使用 8 张 NVIDIA A800。<a target="_blank" rel="noopener noreferrer" class="notion-link" href="https://arxiv.org/pdf/2607.22529#page=4">论文方法与训练细节</a>给出了完整设置。</div><h4 class="notion-h notion-h3 notion-h-indent-1 notion-block-8c5334dff18d400f999b460d8941cfea" data-id="8c5334dff18d400f999b460d8941cfea"><span><div id="8c5334dff18d400f999b460d8941cfea" class="notion-header-anchor"></div><a class="notion-hash-link" href="#8c5334dff18d400f999b460d8941cfea" title="5. 最终推理不带 skill"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">5. 最终推理不带 skill</span></span></h4><div class="notion-text notion-block-81b384c0d23e47379c0962f36ad3c2e3">这是最容易被误读的一点。skill 只帮助构造训练数据和更新模型；最终 solver 仍是普通 prompt-only 模型。论文没有证明运行时加载这些 skill 能带来同样的提升，也没有证明一个现成业务 Agent 接上 skill controller 后就能自行成长。</div><h3 class="notion-h notion-h2 notion-h-indent-0 notion-block-e0459d6ad3b84cd3b8f269fea6ccb64e" data-id="e0459d6ad3b84cd3b8f269fea6ccb64e"><span><div id="e0459d6ad3b84cd3b8f269fea6ccb64e" class="notion-header-anchor"></div><a class="notion-hash-link" href="#e0459d6ad3b84cd3b8f269fea6ccb64e" title="实验结果怎么看"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">实验结果怎么看</span></span></h3><figure class="notion-asset-wrapper notion-asset-wrapper-image notion-block-56a64d8a707b4a96afb8b81827ec6adc"><div style="position:relative;display:flex;justify-content:center;align-self:center;width:100%;max-width:100%;flex-direction:column"><img style="object-fit:cover" src="attachment:09098ea0-775f-4126-b739-46252b465d3b:skill-sp-results.svg?t=56a64d8a-707b-4a96-afb8-b81827ec6adc" alt="notion image" loading="lazy" decoding="async"/></div></figure><div class="notion-text notion-block-cee529e4f3094c258b55869e28e03391"><em>图 2：根据论文 </em><em><a target="_blank" rel="noopener noreferrer" class="notion-link" href="https://arxiv.org/pdf/2607.22529#page=7">Table 1</a></em><em> 与 </em><em><a target="_blank" rel="noopener noreferrer" class="notion-link" href="https://arxiv.org/pdf/2607.22529#page=8">Table 2</a></em><em> 重新绘制。图中是绝对百分点增益，不是相对百分比。</em></div><div class="notion-text notion-block-cd09cee72cc846669cd4979cfa543c3b">论文在两类任务上评测：API-Bank 与 BFCL 测工具选择、参数和格式，ZebraLogic 测约束推理。五个基础模型覆盖 3B 到 14B。</div><div class="notion-text notion-block-b72fcbe01c77498dac7bff3b2f4c278e">结果分成两种情况看更合理。</div><div class="notion-text notion-block-7af623143f974f48a46fb9ccbbffbc52">对本来就会工具调用的 Qwen 与 Granite，工具调用综合分提升 2.8 到 6.5 个点。这部分更像稳定增益。对 Ministral-3-8B 和 14B，工具调用分别提升 42.9 和 42.3 个点，数字很大，但基础分只有 20.7 和 22.2。论文自己把它解释为从严重的 zero-shot schema adherence 失配中恢复。换句话说，Skill-SP 很擅长补齐“知道任务，却不会按工具协议输出”的缺口；这不等于模型的通用推理能力增加了四十多个点。</div><div class="notion-text notion-block-647d68cb05cb42a6aa8b4065babe454d">逻辑推理也全部提升，但弱模型在 Large 和 X-Large 题上几乎没有起色。论文明确承认，自我训练需要最低能力门槛，否则系统无法产生足够可靠的学习信号。Ministral-3-14B 的整体网格准确率增加 12.0 点，主要来自更小的题，而不是最难规模。</div><div class="notion-text notion-block-f5e54199e2fd4ef68ba318b7e519ad31">消融实验比总分更能说明机制：</div><ul class="notion-list notion-list-disc notion-block-38a1ca3475c94d39a75b7ad18d64fe51"><li>去掉 skill 引导后，工具调用综合分从 66.7 降到 64.1；</li></ul><ul class="notion-list notion-list-disc notion-block-7231d61e4b814bd1a070763323199f5e"><li>使用均匀路由降到 64.8；</li></ul><ul class="notion-list notion-list-disc notion-block-69947f3dda2f4befb3ab740a4ac552e4"><li>冻结 skill 库降到 64.4；</li></ul><ul class="notion-list notion-list-disc notion-block-c037c7028076460c9fdb22b224eaee2c"><li>只用 skill stream 也比双通道低 1.2 点。</li></ul><div class="notion-text notion-block-617ef5fe13a14bb7a460d9c1a153cb3a">这些对照支持三个判断：结构化 skill 有用，路由与持续更新有用，自由探索也不能拿掉。它们没有证明每个组件都是全新的。</div><div class="notion-text notion-block-165652c22e874cfeb2e4bb938b56fe13">论文还报告，skill stream 生成的题平均成功率约为 0.57，比自由探索的 0.75 更靠近目标难度；5 轮后有 86 个 active skills，按使用分布折算的 effective skills 为 46。<a target="_blank" rel="noopener noreferrer" class="notion-link" href="https://arxiv.org/pdf/2607.22529#page=10">Figure 5 和 Figure 6</a>说明库没有完全被少数模式占据。这里仍要留个心眼：题目多样性主要靠句向量降维图和使用熵衡量，它们能发现明显聚类，却不能保证语义上没有重复。</div><h3 class="notion-h notion-h2 notion-h-indent-0 notion-block-b10cb067995e44c2b58550cc8f47bf70" data-id="b10cb067995e44c2b58550cc8f47bf70"><span><div id="b10cb067995e44c2b58550cc8f47bf70" class="notion-header-anchor"></div><a class="notion-hash-link" href="#b10cb067995e44c2b58550cc8f47bf70" title="真正的新意在哪里"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">真正的新意在哪里</span></span></h3><div class="notion-text notion-block-8deb3b9482a34f9b8cb40fb66c02296e">我认为最值得带走的不是“模型会自己写 skill”，而是 skill 在训练系统里的职责变了。</div><div class="notion-text notion-block-937f3929990f41c19144e2b867ef38ba">常见 Agent 设计把 skill 放在推理阶段：收到任务后检索一份说明，执行工具，再把结果写进记忆。Skill-SP 把 skill 提前到训练数据生成阶段。一个 skill 同时携带任务结构、局部验证器、难度统计和生命周期状态，因此它可以成为 proposer、solver 与 controller 之间共享的中间层。</div><div class="notion-text notion-block-1573cc36293d47a59eb61170ac005967">这带来三个实际变化：</div><ul class="notion-list notion-list-disc notion-block-3946ebb3f07a4fa096342c9615f5361f"><li>长轨迹不必整段塞回上下文。系统把失败原因和有效模式压缩成可复用 package；</li></ul><ul class="notion-list notion-list-disc notion-block-4b0a323074224155b7804f769ca9c055"><li>验证从全局规则下沉到具体任务模式。不同 skill 可以有不同的检查器；</li></ul><ul class="notion-list notion-list-disc notion-block-e3646a95009d4c6c806957dce7c2be47"><li>curriculum 不再只是一个静态数据集。skill 的成功率、有效性和饱和度决定下一轮出什么题。</li></ul><div class="notion-text notion-block-4a028befda644874b8d2b302cf3db33d">真正新的是这些职责被放进一个共同进化的训练闭环。其余部分多数有成熟来源：proposer-solver 自博弈、GRPO、可验证奖励、中等难度课程、双通道探索、package 式 progressive disclosure 都不是第一次出现。论文的贡献在组合方式和系统位置，不在发明全部零件。</div><h3 class="notion-h notion-h2 notion-h-indent-0 notion-block-5b4864d2941044298b2d1d71e2f961df" data-id="5b4864d2941044298b2d1d71e2f961df"><span><div id="5b4864d2941044298b2d1d71e2f961df" class="notion-header-anchor"></div><a class="notion-hash-link" href="#5b4864d2941044298b2d1d71e2f961df" title="对产品和研发的启发"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">对产品和研发的启发</span></span></h3><div class="notion-text notion-block-91cb657971934e9ebe0d86edb62ce0b3">如果团队想让 Agent 从运行记录中改进，别把“经验”只存成聊天摘要。更实用的沉淀单元应该包含触发条件、执行规则、正反例、自动验收和使用统计。没有验收器的 skill 只能提供建议，很难安全地进入自动优化闭环。</div><div class="notion-text notion-block-ac6a8ed276154bbd87fea09cec9e3c60">课程也不该只追求更多数据。让当前 Agent 反复失败的超难任务没有训练价值，已经稳定通过的任务也不值得占预算。可以用历史成功率建立一个能力边界池，把训练和评测资源集中在“有机会学会，但尚未稳定”的任务上。</div><div class="notion-text notion-block-02a38af5863b40eb8bead59984c21274">探索与复用要分开记账。稳定 skill 负责高质量产出，开放通道负责找新模式；新模式经过验证后再升格为 skill。这样比让一个生成器同时兼顾可靠与新颖更容易治理。</div><div class="notion-text notion-block-a76ad1b677b543068b0ae0eaef1e8293">对于普通应用团队，这套方法目前更适合借鉴结构，而不是原样复刻训练。8 张 A800、双模型滚动生成与多次 probe 的成本不低。更轻的做法是先在运行时收集轨迹，用确定性测试维护 skill 库，再定期离线训练或人工审核。</div><h3 class="notion-h notion-h2 notion-h-indent-0 notion-block-e747466b718a4617bc4187d8e04257fb" data-id="e747466b718a4617bc4187d8e04257fb"><span><div id="e747466b718a4617bc4187d8e04257fb" class="notion-header-anchor"></div><a class="notion-hash-link" href="#e747466b718a4617bc4187d8e04257fb" title="风险、局限和尚未验证的问题"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">风险、局限和尚未验证的问题</span></span></h3><figure class="notion-asset-wrapper notion-asset-wrapper-image notion-block-d21d4849bbf74f64935a800c2ad14262"><div style="position:relative;display:flex;justify-content:center;align-self:center;width:100%;max-width:100%;flex-direction:column"><img style="object-fit:cover" src="attachment:a58ca89b-5979-4eaf-ba5f-70fda0f054a8:skill-sp-boundary.svg?t=d21d4849-bbf7-4f64-935a-800c2ad14262" alt="notion image" loading="lazy" decoding="async"/></div></figure><div class="notion-text notion-block-553eda9366d74392a8f58d41313745d8"><em>图 3：根据论文实验范围与 Limitations 章节整理。绿色部分有公开实验支撑，橙色部分还没有。来源：</em><em><a target="_blank" rel="noopener noreferrer" class="notion-link" href="https://arxiv.org/pdf/2607.22529#page=22">论文附录 H</a></em><em>。</em></div><div class="notion-text notion-block-309d0647b3264ada9115ad89b017e211">第一，实验只有工具调用预测和逻辑网格题。它们都能写出确定性验收器，和长周期浏览器、computer use、多人协作、状态恢复、人机确认不是一类难题。论文标题里的“LLM capability frontier”说得比证据范围更宽。</div><div class="notion-text notion-block-e5d1e0aee3994028858c1c937442a4c5">第二，skill induction 仍依赖基础模型具备最低能力。作者也写明，复杂领域可能要用少量人工示范启动。所谓“开放式自我进化”目前不是从零开始。</div><div class="notion-text notion-block-5bd03181191a41b99fc03eb98b5af8e7">第三，路由有固定参数。skill 与探索数据的混合比例、难度区间和归档门槛需要为新任务重新调。自动生成规则与 validator 也被列为未来工作，说明当前验证边界仍有人工设计成分。</div><div class="notion-text notion-block-f17c0795db3645d989cb8324fe4548cb">第四，语义新颖性检查使用词面相似度阈值。它可能保留换了说法的重复 skill，也可能误删表述接近但约束不同的 skill。validator 一旦写错，错误会比普通提示词更隐蔽，因为它会直接决定训练奖励。</div><div class="notion-text notion-block-333806df8f8d4efea4e14070d37568f6">第五，训练只观察 5 轮，没有报告更长周期下的库膨胀、遗忘、退化、训练耗时或成本。跨模型迁移 skill 库也只是未来方向。</div><div class="notion-text notion-block-5e38d79c3a094ec9b0fc8da9fdc3fea7">第六，开源仓库是一次性首发提交。它包含训练脚本、评测数据和 15 个初始工具调用 package，README 假设 8 张 GPU。我拉取仓库做了静态检查，Python 文件能通过语法编译；但仓库没有自动化测试目录，也没有论文训练日志或现成 checkpoint。我没有条件复跑 8×A800 实验，因此正文中的分数仍属于作者报告，不能写成独立复现结果。</div><h3 class="notion-h notion-h2 notion-h-indent-0 notion-block-68fe0303ad374d95b70168cb5d277d0f" data-id="68fe0303ad374d95b70168cb5d277d0f"><span><div id="68fe0303ad374d95b70168cb5d277d0f" class="notion-header-anchor"></div><a class="notion-hash-link" href="#68fe0303ad374d95b70168cb5d277d0f" title="今日沉淀"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">今日沉淀</span></span></h3><ol start="1" class="notion-list notion-list-numbered notion-block-2c6a40beee2b408aa9441b62f4395f8a"><li>能进入自我优化闭环的 skill，至少要有触发条件、规则、validator 和使用统计。</li></ol><ol start="2" class="notion-list notion-list-numbered notion-block-73df917df1414d5592fbe470d6792979"><li>自博弈需要两条数据线：结构化复用保证质量，开放探索负责发现新模式。</li></ol><ol start="3" class="notion-list notion-list-numbered notion-block-56157a4dd443406498c83f7767a7389c"><li>好课程不追最难题，而追当前成功率接近一半的能力边界。</li></ol><ol start="4" class="notion-list notion-list-numbered notion-block-207da84141084ebea7c77c3a5054fd8a"><li>训练时 skill 与推理时 skill 是两种产品，不能把结论混用。</li></ol><ol start="5" class="notion-list notion-list-numbered notion-block-8ea45837d1ed42ef82aaea309eba6f42"><li>弱模型的大幅增益先检查是否只是协议与格式对齐，不要直接解释成通用智能跃升。</li></ol></main></div>]]></content:encoded>
        </item>
        <item>
            <title><![CDATA[Agentic Context Management 深度分析 - 2026-07-27]]></title>
            <link>http://easyai.fyi/article/agentic-context-management-deep-analysis-2026-07-27</link>
            <guid>http://easyai.fyi/article/agentic-context-management-deep-analysis-2026-07-27</guid>
            <pubDate>Mon, 27 Jul 2026 00:00:00 GMT</pubDate>
            <description><![CDATA[Agentic Context Management 把 Agent 记忆从存取功能改写成覆盖架构、摄取、作用域、预取和可验证压缩的完整生命周期。]]></description>
            <content:encoded><![CDATA[<div id="notion-article" class="mx-auto overflow-hidden "><main class="notion light-mode notion-page notion-block-3aa5eac3813c8149ab90eef91647dc88"><div class="notion-viewport"></div><div class="notion-collection-page-properties"></div><h3 class="notion-h notion-h2 notion-h-indent-0 notion-block-2b66cc1218284641b47d34d344c73a02" data-id="2b66cc1218284641b47d34d344c73a02"><span><div id="2b66cc1218284641b47d34d344c73a02" class="notion-header-anchor"></div><a class="notion-hash-link" href="#2b66cc1218284641b47d34d344c73a02" title="原始链接"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">原始链接</span></span></h3><ul class="notion-list notion-list-disc notion-block-b8eed8f5c6cd44fcab14b54dfa1719f0"><li>论文页：<a target="_blank" rel="noopener noreferrer" class="notion-link" href="https://arxiv.org/abs/2607.21503">Agentic Context Management</a></li></ul><ul class="notion-list notion-list-disc notion-block-70292c49375f4ae4a9be1f7b7cf3c2e1"><li>论文 HTML：<a target="_blank" rel="noopener noreferrer" class="notion-link" href="https://arxiv.org/html/2607.21503v1">arXiv HTML 全文</a></li></ul><ul class="notion-list notion-list-disc notion-block-5391b2d82994413ea46434a82d5c9fca"><li>论文 PDF：<a target="_blank" rel="noopener noreferrer" class="notion-link" href="https://arxiv.org/pdf/2607.21503">arXiv PDF</a></li></ul><ul class="notion-list notion-list-disc notion-block-ff295e80aa50488f9d81d029b90cfe01"><li>Hugging Face Papers：<a target="_blank" rel="noopener noreferrer" class="notion-link" href="https://huggingface.co/papers/2607.21503">论文社区页</a></li></ul><ul class="notion-list notion-list-disc notion-block-48864de8ec61400b81076596539159ac"><li>参考实现 SDK：<a target="_blank" rel="noopener noreferrer" class="notion-link" href="https://github.com/maximem-ai/maximem_synap_sdk">maximem-ai/maximem_synap_sdk</a></li></ul><ul class="notion-list notion-list-disc notion-block-720deee9bdf248b39c7d2f276dfc9f73"><li>评测工具：<a target="_blank" rel="noopener noreferrer" class="notion-link" href="https://github.com/maximem-ai/memory_and_context_eval_harness">memory_and_context_eval_harness</a></li></ul><ul class="notion-list notion-list-disc notion-block-0aa9e167bd36492e81b66b675a9255d3"><li>评测结果记录：<a target="_blank" rel="noopener noreferrer" class="notion-link" href="https://github.com/maximem-ai/eval_benchmark_runs_output">eval_benchmark_runs_output</a></li></ul><ul class="notion-list notion-list-disc notion-block-6168101812ec4ba0a294658700b3ee32"><li>项目文档：<a target="_blank" rel="noopener noreferrer" class="notion-link" href="https://docs.maximem.ai">Maximem Synap Docs</a></li></ul><h3 class="notion-h notion-h2 notion-h-indent-0 notion-block-96bc87eadf2542f1b7e2747f307d18b6" data-id="96bc87eadf2542f1b7e2747f307d18b6"><span><div id="96bc87eadf2542f1b7e2747f307d18b6" class="notion-header-anchor"></div><a class="notion-hash-link" href="#96bc87eadf2542f1b7e2747f307d18b6" title="为什么今天选它"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">为什么今天选它</span></span></h3><div class="notion-text notion-block-95ca4b9791a8454fb8745a213c4593b7">今天重点比较了三个对象：<a target="_blank" rel="noopener noreferrer" class="notion-link" href="https://arxiv.org/abs/2607.21503">Agentic Context Management</a> 把 Agent 记忆重写成一套上下文生命周期；<a target="_blank" rel="noopener noreferrer" class="notion-link" href="https://arxiv.org/abs/2607.11388">StructAgent</a> 用可验证的因果状态推进长周期电脑操作；<a target="_blank" rel="noopener noreferrer" class="notion-link" href="https://arxiv.org/abs/2605.04808">DTap</a> 提供可控的 Agent 红队环境。</div><div class="notion-text notion-block-9f730156c8f242c9a65567dbaeed7f49">StructAgent 的结果很强，但【智汇AI】最近已经写过 ScaleCUA、SearchOS 和长周期状态管理。DTap 的安全评测也与 Mako、Safety Sentry 相邻，而且论文最初在 5 月提交。Agentic Context Management 于 2026 年 7 月 23 日提交，7 月 27 日进入 Hugging Face Daily Papers，当天排第 3；我检查时有 15 个赞，配套 SDK 仓库有 50 stars。热度还不算大，但材料完整，问题也更贴近生产：Agent 到底该在每一步“记住什么”，而不是数据库里“存了什么”。<a target="_blank" rel="noopener noreferrer" class="notion-link" href="https://arxiv.org/abs/2607.21503">提交记录</a> <a target="_blank" rel="noopener noreferrer" class="notion-link" href="https://huggingface.co/papers/2607.21503">Daily Papers 页面</a> <a target="_blank" rel="noopener noreferrer" class="notion-link" href="https://github.com/maximem-ai/maximem_synap_sdk">GitHub 仓库</a></div><h3 class="notion-h notion-h2 notion-h-indent-0 notion-block-f0e98aca6ece46099c99844e68fae4bd" data-id="f0e98aca6ece46099c99844e68fae4bd"><span><div id="f0e98aca6ece46099c99844e68fae4bd" class="notion-header-anchor"></div><a class="notion-hash-link" href="#f0e98aca6ece46099c99844e68fae4bd" title="先说判断：它改的不是存储，而是决策边界"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">先说判断：它改的不是存储，而是决策边界</span></span></h3><div class="notion-text notion-block-677a438dd7ff4d628786aecf4533e539">多数 Agent memory 产品围绕两个动作设计：写入、召回。论文认为这还不够。一个真正跑在生产里的 Agent，每一轮都要回答更多问题：刚发生的内容值不值得留；应该提取成什么结构；属于个人、公司还是平台；现在该拿出哪一小部分；下一步可能缺什么；窗口装不下时，删掉什么才不会破坏后续推理。</div><div class="notion-text notion-block-ddf20d24e9a947feb43025e30ccd4878">这组判断不能交给向量库。向量库只负责找“相似的东西”，不会替系统决定保留期限、权限边界、上下文预算和信息损失。论文因此提出 Agentic Context Management（ACM），把上下文管理拆成五个环节：<a target="_blank" rel="noopener noreferrer" class="notion-link" href="https://arxiv.org/html/2607.21503v1#S2">论文第 2 节</a></div><ol start="1" class="notion-list notion-list-numbered notion-block-90eba61b18db4068ad38bd84d6cd34b1"><li>Architecting：先按 Agent 的用途设计记忆类型、保留规则、检索和压缩方式。</li></ol><ol start="2" class="notion-list notion-list-numbered notion-block-c1918e5b594a468183ba40c4fa66917b"><li>Ingesting：把对话、文档、工具调用和返回结果提取成可检索的事实、关系、事件与时间信息。</li></ol><ol start="3" class="notion-list notion-list-numbered notion-block-b42e49f1f83d4f89b75c2334f46ef873"><li>Scoping：决定内容属于用户、客户组织还是平台，并在写入和读取时保持隔离。</li></ol><ol start="4" class="notion-list notion-list-numbered notion-block-8db1b2c06e45415395a95467f00dc3ca"><li>Anticipating：不等 Agent 发起查询，先预测下一步可能需要的上下文并预取。</li></ol><ol start="5" class="notion-list notion-list-numbered notion-block-9e40b1d67b49492eb22f830f47cb8a6c"><li>Compacting &amp; Consolidation：把上下文压进预算，同时检查关键事实是否仍能恢复。</li></ol><figure class="notion-asset-wrapper notion-asset-wrapper-image notion-block-22686621e3044c22aaace5391f705d6b"><div style="position:relative;display:flex;justify-content:center;align-self:center;width:100%;max-width:100%;flex-direction:column"><img style="object-fit:cover" src="https://arxiv.org/html/2607.21503v1/x1.png?t=22686621-e304-4c22-aaac-e5391f705d6b" alt="图 1：ACM 的五环生命周期与组织作用域。" loading="lazy" decoding="async"/><figcaption class="notion-asset-caption">图 1：ACM 的五环生命周期与组织作用域。</figcaption></div></figure><div class="notion-text notion-block-8e9cfd110de74eb683f0e5a94fb87eca">图里真正重要的是两条轴同时存在。横向是上下文从设计到退休的生命周期，纵向是 user、customer、client 三层作用域。作者给出的边界很明确：只做好其中一个环节，仍然是 memory tool；五个环节能协同工作，并覆盖组织作用域，才算 context-management platform。<a target="_blank" rel="noopener noreferrer" class="notion-link" href="https://arxiv.org/html/2607.21503v1#S2.F1">来源：论文 Figure 1</a></div><h3 class="notion-h notion-h2 notion-h-indent-0 notion-block-92282067297b4ac1a8624b0d371a06ac" data-id="92282067297b4ac1a8624b0d371a06ac"><span><div id="92282067297b4ac1a8624b0d371a06ac" class="notion-header-anchor"></div><a class="notion-hash-link" href="#92282067297b4ac1a8624b0d371a06ac" title="系统怎么运作"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">系统怎么运作</span></span></h3><div class="notion-text notion-block-ae233eb72e0d4ded8960f5c60dfc1fdd">论文用 Maximem Synap 作为参考实现。接入一个新 Agent 时，系统先根据用途说明和参考材料生成一份专用 memory architecture。客服 Agent、编程 Agent 和语音 Agent 的记忆类型、保留时间、压缩规则不同，不再共用固定 schema。作者称这个设计步骤由 LLM 完成，并带多 Agent 检查，但具体生成和验证机制没有公开。<a target="_blank" rel="noopener noreferrer" class="notion-link" href="https://arxiv.org/html/2607.21503v1#S4">系统结构</a></div><div class="notion-text notion-block-2d84a59c34f04f619e623aa6abe4c67a">新对话、文件和工具结果进入后，写入走异步管线。调用方先拿到 ingestion ID，实体识别、关系提取和多种存储写入在后台完成。为了避免“刚说完的话下一轮还没写好”，最近几轮原文继续留在工作上下文里；异步处理只影响长期结构化记忆，不影响同一会话的即时读取。</div><div class="notion-text notion-block-a76df0fa2c344ddfbee56d46f6b2386d">检索时，系统按 user → customer → client 的顺序由窄到宽地找信息，每条结果带来源，并被裁进 token 预算。它把向量相似度和图关系结合起来：向量负责找到语义入口，图负责补上多跳关系。单靠相似度经常能找到一个相关文档，却漏掉完成推理所需的“桥接文档”。</div><div class="notion-text notion-block-a8bc403a29fc4429bc2caa6c7f210243">Anticipating 是这篇论文最值得单独留下的机制。普通 retrieval 回答“现在这条查询相关什么”，它试图回答“Agent 下一步会需要什么”。预测命中后，真正使用时只读预取结果，不再把检索延迟放在 Agent 的主路径上。论文称跨客户能稳定达到 60% 以上命中率，但没有公开预测方法、客户分布和浪费的预取成本，因此这还是产品报告数字，不是可复现结论。<a target="_blank" rel="noopener noreferrer" class="notion-link" href="https://arxiv.org/html/2607.21503v1#S4">预取说明</a></div><div class="notion-text notion-block-d43c16fad0304614aa59a3d04300d9c3">长对话接近预算时，系统做 category-aware compaction。需要原样保留的字段不动，允许抽象的内容才压缩；压缩后再检查原文中的关键信息能否恢复，低于阈值就降低压缩力度重试。每次返回 validation score 和 compression ratio。这个设计比“每隔十轮总结一次”多了一条质量契约：压缩必须交代自己丢没丢东西。</div><figure class="notion-asset-wrapper notion-asset-wrapper-image notion-block-648a26a0c7ae4df8b0f6b76e7ebf7135"><div style="position:relative;display:flex;justify-content:center;align-self:center;width:100%;max-width:100%;flex-direction:column"><img style="object-fit:cover" src="https://arxiv.org/html/2607.21503v1/x5.png?t=648a26a0-c7ae-4df8-b0f6-b76e7ebf7135" alt="图 2：Synap 参考架构。SDK 经 API 接入，五个上下文环节由不同管线实现，底层组合关系、向量、对象和短期状态存储。" loading="lazy" decoding="async"/><figcaption class="notion-asset-caption">图 2：Synap 参考架构。SDK 经 API 接入，五个上下文环节由不同管线实现，底层组合关系、向量、对象和短期状态存储。</figcaption></div></figure><div class="notion-text notion-block-b49c85c1b40f4560abbd56971883b118">应用侧最终只需要围绕一次模型调用做三件事：取回过去上下文，必要时压缩当前对话，再把新一轮内容异步写入。图中的队列、多种数据库、缓存和 API 并不新；新意在于它们都服从同一份 per-agent architecture、同一套作用域和同一个上下文预算。<a target="_blank" rel="noopener noreferrer" class="notion-link" href="https://arxiv.org/html/2607.21503v1#S4.F5">来源：论文 Figure 5</a></div><h3 class="notion-h notion-h2 notion-h-indent-0 notion-block-fa525369397242d788f8a2d36540cd0b" data-id="fa525369397242d788f8a2d36540cd0b"><span><div id="fa525369397242d788f8a2d36540cd0b" class="notion-header-anchor"></div><a class="notion-hash-link" href="#fa525369397242d788f8a2d36540cd0b" title="为什么“把全部历史塞进去”迟早会出问题"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">为什么“把全部历史塞进去”迟早会出问题</span></span></h3><div class="notion-text notion-block-3f4372c5665347dd89ddc3856afc43ba">如果每一轮新增约 t 个 token，第 k 轮又把前 k 轮全部重发，n 轮累计输入量是 O(n²)。论文用 t=500、固定窗口预算 W=4000 举例：100 轮时，全量追加的累计输入约为固定预算的 6 倍；200 轮时约为 13 倍。这个推导算的是累计输入账单，不是某个模型的实测延迟。<a target="_blank" rel="noopener noreferrer" class="notion-link" href="https://arxiv.org/html/2607.21503v1#S3.SS1">成本推导</a></div><figure class="notion-asset-wrapper notion-asset-wrapper-image notion-block-224b084eef4d4843a066cf18425e18f1"><div style="position:relative;display:flex;justify-content:center;align-self:center;width:100%;max-width:100%;flex-direction:column"><img style="object-fit:cover" src="https://arxiv.org/html/2607.21503v1/x2.png?t=224b084e-ef4d-4843-a066-cf18425e18f1" alt="图 3：全量追加与固定预算的累计输入量。" loading="lazy" decoding="async"/><figcaption class="notion-asset-caption">图 3：全量追加与固定预算的累计输入量。</figcaption></div></figure><div class="notion-text notion-block-8b98aa4e38584ba7b2a16c0a39c10b71">这张图说明了为什么“模型窗口再大一点”治不了成本增长。窗口变大只是把爆点推迟。系统仍需要主动决定每轮放什么，而不是任由历史线性变长、累计费用二次增长。<a target="_blank" rel="noopener noreferrer" class="notion-link" href="https://arxiv.org/html/2607.21503v1#S3.F2">来源：论文 Figure 2</a></div><div class="notion-text notion-block-e0e00b57fe6b41368f70f61b3b3345a8">但简单摘要也不可靠。论文引用 ACE 的一个案例：18,282 tokens 被一次压到 122 tokens 后，任务准确率从 66.7% 降到 57.1%，甚至低于无上下文基线。ACM 因此追求“线性成本 + 已检查的信息保真”。这里要注意，validated compaction 是目标设计，不等于论文已经证明它在各种任务上无损。<a target="_blank" rel="noopener noreferrer" class="notion-link" href="https://arxiv.org/html/2607.21503v1#S3.SS2">压缩讨论</a></div><div class="notion-text notion-block-d682b7a523e24613955b7ed26b412c86">按论文给出的另一个示例，窗口维持在 4000 tokens、每 8 轮压缩一次、每次验证成本相当于两倍窗口，验证会带来固定 25% 开销；与全量追加相比，净节省会随对话拉长，在 100、200、500 轮时约为 80%、90%、96%。这些仍是公式推演，不是线上账单。<a target="_blank" rel="noopener noreferrer" class="notion-link" href="https://arxiv.org/html/2607.21503v1#S5">设计选择与成本</a></div><h3 class="notion-h notion-h2 notion-h-indent-0 notion-block-f1191b3168814380af0e4877d5b56d96" data-id="f1191b3168814380af0e4877d5b56d96"><span><div id="f1191b3168814380af0e4877d5b56d96" class="notion-header-anchor"></div><a class="notion-hash-link" href="#f1191b3168814380af0e4877d5b56d96" title="检索命中不等于上下文够用"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">检索命中不等于上下文够用</span></span></h3><div class="notion-text notion-block-598993180f6c4b65816fc449598a6ade">论文把答案质量的上限写成三个瓶颈的最小值：提取质量、检索质量、推理所需信息是否齐全。写入时把“用户 4 月 3 日从 Starter 升到 Pro”压成“用户提到过套餐”，之后用再好的检索也找不回日期和变化。找到一篇相关材料，却没找到完成多跳推理所需的第二篇，也只算“命中”，不能算“够用”。<a target="_blank" rel="noopener noreferrer" class="notion-link" href="https://arxiv.org/html/2607.21503v1#S3.SS3">检索充分性</a></div><div class="notion-text notion-block-09e61d53dcb14e37a46d63acc1354e04">作者做了一个动机性实验：CodeXGLUE、MS MARCO、SQuAD、HotpotQA、SciQ 五个数据集，各取 10,000 篇文档和 1,000 条查询，对比关键词检索与向量检索。自然语言找代码时，向量 MRR@10 为 0.914，关键词为 0.290；SciQ 科学问答里，关键词为 0.815，向量为 0.614。向量索引同一批 10,000 篇文档需要多 60 到 100 倍时间。<a target="_blank" rel="noopener noreferrer" class="notion-link" href="https://arxiv.org/html/2607.21503v1#A2">实验方法与数字</a></div><figure class="notion-asset-wrapper notion-asset-wrapper-image notion-block-0c21ad344f36401c9dca659907be83df"><div style="position:relative;display:flex;justify-content:center;align-self:center;width:100%;max-width:100%;flex-direction:column"><img style="object-fit:cover" src="https://arxiv.org/html/2607.21503v1/x4.png?t=0c21ad34-4f36-401c-9dca-659907be83df" alt="图 4：五类数据上的关键词与向量检索结果，以及向量索引成本。" loading="lazy" decoding="async"/><figcaption class="notion-asset-caption">图 4：五类数据上的关键词与向量检索结果，以及向量索引成本。</figcaption></div></figure><div class="notion-text notion-block-234c47fa617b4e1ba2f80ee6f91c885b">这组结果支持“不同查询需要不同信号”，不支持“关键词一定优于向量”或反过来。实验只有一套关键词引擎、一套向量库、一个 embedding 模型，没有 chunking，单库规模也只有 10,000 篇；长文还会被 embedding 模型截断。作者把它定位为动机实验是合适的。<a target="_blank" rel="noopener noreferrer" class="notion-link" href="https://arxiv.org/html/2607.21503v1#A2">来源：论文 Figure 4 与 Appendix B</a></div><div class="notion-text notion-block-0411032ff011416987bf212fc539783c">更有价值的是它暴露了评测盲点：HotpotQA 实际需要多篇支持文档，但这次只按一个 gold target 计分。这样的 MRR 只能判断“有没有碰到相关材料”，无法判断“信息是否足以完成推理”。生产 Agent 应同时测命中、缺口、冗余和最终任务结果。</div><h3 class="notion-h notion-h2 notion-h-indent-0 notion-block-34889fbd79644fb69bed22ae1a719ae8" data-id="34889fbd79644fb69bed22ae1a719ae8"><span><div id="34889fbd79644fb69bed22ae1a719ae8" class="notion-header-anchor"></div><a class="notion-hash-link" href="#34889fbd79644fb69bed22ae1a719ae8" title="结果怎么读：92% 很高，但证据还不完整"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">结果怎么读：92% 很高，但证据还不完整</span></span></h3><div class="notion-text notion-block-4ada3db3abeb4e089d25a0c78dbfb525">Synap 在 LongMemEval 的 500 道题上答对 460 道，得到 92.0%；在 LoCoMo 第 1—4 类上报告 93.2%。回答模型和裁判模型都是 gpt-5-mini。LoCoMo 的第 5 类是不可回答问题，测的是拒答而不是记忆，论文按原论文、Mem0 和 Zep 的常见口径将其排除。作者也提醒，不同系统使用的回答模型、裁判、摄取粒度和类别口径不同，不能把各家自报分数直接排成排行榜。<a target="_blank" rel="noopener noreferrer" class="notion-link" href="https://arxiv.org/html/2607.21503v1#S6">评测设置</a></div><figure class="notion-asset-wrapper notion-asset-wrapper-image notion-block-2f6651e7f94544d6841609792783229c"><div style="position:relative;display:flex;justify-content:center;align-self:center;width:100%;max-width:100%;flex-direction:column"><img style="object-fit:cover" src="https://arxiv.org/html/2607.21503v1/x6.png?t=2f6651e7-f945-44d6-8416-09792783229c" alt="图 5：LongMemEval 与 LoCoMo 的分类结果。" loading="lazy" decoding="async"/><figcaption class="notion-asset-caption">图 5：LongMemEval 与 LoCoMo 的分类结果。</figcaption></div></figure><div class="notion-text notion-block-a34067f35d1f433187e7ce7402e28b46">分类结果比总分更有用。LongMemEval 的单会话用户信息、偏好、知识更新和时间推理都达到 100%，但跨会话推理只有 75.2%，是最明显的短板。它说明“能记住单条事实”和“能把多次会话的信息拼成结论”仍是两种能力。<a target="_blank" rel="noopener noreferrer" class="notion-link" href="https://arxiv.org/html/2607.21503v1#S6.F6">来源：论文 Figure 6</a></div><div class="notion-text notion-block-9824969132c84b29b236c845b732454e">这项评测没有测生产负载下的延迟、单任务 token 成本、上下文变长后的抗干扰，也没有测真实工具调用和多 Agent 交接。论文没有给出线上延迟数字，完整逐题输出目前只能向作者索取；公开的是评测工具、配置和分类统计。核心引擎也没有开源，GitHub 仓库只包含 Python、JavaScript SDK、框架适配和 MCP 适配器，使用时必须连接 Maximem 的托管服务。<a target="_blank" rel="noopener noreferrer" class="notion-link" href="https://arxiv.org/html/2607.21503v1#S6.SS3">论文限制</a> <a target="_blank" rel="noopener noreferrer" class="notion-link" href="https://github.com/maximem-ai/maximem_synap_sdk#the-memory-layer-for-production-ai-agents">SDK README</a></div><div class="notion-text notion-block-617f6f19c4a64242b192058a8bcd059a">所以，92% 应读作“作者在公开记忆测试集和已说明配置上取得的结果”，不能读成“这套上下文生命周期已经被第三方验证为生产最优”。</div><h3 class="notion-h notion-h2 notion-h-indent-0 notion-block-2ee267197caf47e9acd4532ce1980c01" data-id="2ee267197caf47e9acd4532ce1980c01"><span><div id="2ee267197caf47e9acd4532ce1980c01" class="notion-header-anchor"></div><a class="notion-hash-link" href="#2ee267197caf47e9acd4532ce1980c01" title="真正新的地方，哪些只是工程包装"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">真正新的地方，哪些只是工程包装</span></span></h3><div class="notion-text notion-block-7096968537e2433fa1fb5bc837c7c8aa">“外部记忆”“上下文分层”“自动压缩”都不是新想法。MemGPT / Letta 已经把 LLM 类比成操作系统，让模型在工作窗口与外部记忆之间换页；GraphRAG、HippoRAG、Zep / Graphiti 已经把关系图用于多跳检索；异步写入、混合存储、实体归一和多租户隔离也是成熟的软件设计。<a target="_blank" rel="noopener noreferrer" class="notion-link" href="https://arxiv.org/html/2607.21503v1#S7">相关工作</a></div><div class="notion-text notion-block-18b8fd2a58f44ffe8d661e74ae5fe496">这篇论文新增的主要是四个明确边界：</div><ul class="notion-list notion-list-disc notion-block-d5a4e754032448068d050eb1a18890f5"><li>先为每类 Agent 生成 memory architecture，再让摄取、检索、压缩都服从它。</li></ul><ul class="notion-list notion-list-disc notion-block-5fd9fd97cdcc4beeb4fdca30d04eb7d7"><li>把组织作用域放进写入和读取协议，而不是只在提示词里提醒“不要串数据”。</li></ul><ul class="notion-list notion-list-disc notion-block-607901413e5b4e61ac790e9825578ad0"><li>把预测性预取视为独立环节，目标是把 retrieval 从主路径移走。</li></ul><ul class="notion-list notion-list-disc notion-block-61ddf35c03e44c1a87324116b797fc1c"><li>把压缩当作需要验收的状态变换，输出可检查分数，失败就重试。</li></ul><div class="notion-text notion-block-c87637f070d64c2fae604c18dfc825c0">这四点放在一起，确实比“接一个向量库，再写段 summary prompt”完整。论文最好的贡献是统一问题定义，而不是某个新算法。</div><div class="notion-text notion-block-a4c1ceb9e5124c64b5ed78b566f0481b">工程包装也很明显。参考实现来自作者所在公司，只有一位作者，几处最关键机制——memory architecture 如何生成、预取如何预测、压缩如何验证、图遍历如何打分——都以商业机密为由省略。系统覆盖五个环节的对比表主要依据各家公开文档，由作者自行定义口径。它更像一篇带产品实例的类别论文，而不是把每个模块都拆开验证的学术系统论文。</div><h3 class="notion-h notion-h2 notion-h-indent-0 notion-block-5f855e69300c41dd8107312b303eada3" data-id="5f855e69300c41dd8107312b303eada3"><span><div id="5f855e69300c41dd8107312b303eada3" class="notion-header-anchor"></div><a class="notion-hash-link" href="#5f855e69300c41dd8107312b303eada3" title="对产品和研发的启发"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">对产品和研发的启发</span></span></h3><div class="notion-text notion-block-433f321c894f46d383c3411ca0c96954">第一步不是选向量库，而是做上下文清单。对每类信息写清楚来源、结构、有效期、敏感级别、归属范围、召回条件和允许的压缩方式。没有这张表，memory 会很快变成无边界的数据堆。</div><div class="notion-text notion-block-bf59e8af33ba439b8aa82f1dd2078bcf">把最近上下文和长期上下文分开。新一轮消息先原样留在工作窗口，保证 read-your-writes；长期提取、实体合并和持久化可以异步完成。这样能减少延迟，又不会出现“用户刚说完，Agent 下一轮就忘了”。</div><div class="notion-text notion-block-f0e60021874242d1acd6412f17d05c0d">作用域要落在存储与查询层。user、team、organization、platform 需要独立身份和明确的合并顺序。只靠 prompt 约束跨租户读取，迟早会出事。</div><div class="notion-text notion-block-e56bfa41e186428cb5ffe12f694b0685">给 compaction 建回归测试。准备一批必须保留的事实、约束、未完成事项和决策理由。每次修改摘要规则、模型或预算，都检查这些信息压缩后还能否恢复。压缩比不能单独作为目标。</div><div class="notion-text notion-block-f6bdd091adde4232a37ac3654f97ed5e">预取先跑影子模式。系统可以预测下一步所需内容，但先不影响正式上下文，只记录命中率、额外算力、延迟收益和误取敏感数据的比例。真正省下主路径时间后，再逐步启用。</div><div class="notion-text notion-block-4ef7a3aaa74745f1a63150102b70877c">评测要从 retrieval hit 升到 reasoning sufficiency。除了 Top-k 和 MRR，还要问：完成任务所需证据是否齐全；无关内容占了多少窗口；来源是否可追；缺失时 Agent 会不会停下来补取；跨会话和跨 Agent 交接后能否继续工作。</div><h3 class="notion-h notion-h2 notion-h-indent-0 notion-block-d251534d67804595abe4d6c865b02b0b" data-id="d251534d67804595abe4d6c865b02b0b"><span><div id="d251534d67804595abe4d6c865b02b0b" class="notion-header-anchor"></div><a class="notion-hash-link" href="#d251534d67804595abe4d6c865b02b0b" title="风险、局限和尚未验证的问题"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">风险、局限和尚未验证的问题</span></span></h3><div class="notion-text notion-block-c336fa9508294e31b0a8a6674624c9a9">最大问题是可复现性。公开 SDK 能证明接口和适配范围，不能检查托管引擎是否真的按论文描述工作。逐题结果没有直接公开，预取、压缩验证和自动架构生成也缺少可复现实验。</div><div class="notion-text notion-block-7497c25364a342d5add45b9eebcde4d8">评测范围偏窄。LongMemEval 和 LoCoMo 主要测试对话记忆，不能代表编程 Agent 的仓库状态、浏览器 Agent 的页面变化、企业 Agent 的权限与审批链。论文提出的“组织级上下文”恰好没有被这两套评测覆盖。</div><div class="notion-text notion-block-20935170f53c4d17845675a525795a18">同一个 gpt-5-mini 同时回答和裁判，可能带来一致性偏差。LoCoMo 排除拒答类问题也让总分更高，虽然作者清楚说明了口径。更稳的验证需要独立裁判、人工抽检和含不可回答问题的完整结果。</div><div class="notion-text notion-block-0517b59322ee4773b8889cfc93d8a591">自动生成 memory architecture 可能不稳定。若它把敏感字段放到过宽作用域，或给关键事实设置了过短有效期，后续每个环节都会继承这个错误。生产系统需要可审查的配置、版本记录、回滚和人工批准，不能把架构生成当成一次性提示词任务。</div><div class="notion-text notion-block-97f42525bc614b929fe02b90181732c0">“主动遗忘”还牵涉审计与合规。用户要求删除、业务要求保留、模型需要压缩，这三件事不是同一个动作。物理删除、逻辑失效、摘要合并和降低检索权重必须分开记录，否则系统无法解释某条信息为什么消失。</div><h3 class="notion-h notion-h2 notion-h-indent-0 notion-block-a996fdad3ab7439cb23a1ee2a4d2f319" data-id="a996fdad3ab7439cb23a1ee2a4d2f319"><span><div id="a996fdad3ab7439cb23a1ee2a4d2f319" class="notion-header-anchor"></div><a class="notion-hash-link" href="#a996fdad3ab7439cb23a1ee2a4d2f319" title="今日沉淀"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">今日沉淀</span></span></h3><ol start="1" class="notion-list notion-list-numbered notion-block-89cd4b7545fc43c6a8758921e8790c87"><li>Agent memory 的核心不是存得多，而是每一步该带什么进入上下文。</li></ol><ol start="2" class="notion-list notion-list-numbered notion-block-071880dac0244d3cb8ddf122c9ff102a"><li>写入、作用域、预取和压缩必须共用一份上下文规则。</li></ol><ol start="3" class="notion-list notion-list-numbered notion-block-ab88cb1a869448de9588c8533b9e74be"><li>压缩要验收信息是否还在，不能只看省了多少 token。</li></ol><ol start="4" class="notion-list notion-list-numbered notion-block-ead64efed0b14028b1537b158470ec3c"><li>检索命中不等于证据齐全，评测要看能否完成推理。</li></ol><ol start="5" class="notion-list notion-list-numbered notion-block-e73f64dda9cf40a48022c1c975993735"><li>组织级记忆先解决隔离和审计，再谈共享带来的增益。</li></ol></main></div>]]></content:encoded>
        </item>
        <item>
            <title><![CDATA[Agent 工作流和推理基建同时加速 - 2026-07-24]]></title>
            <link>http://easyai.fyi/article/follow-builders-ai-summary-2026-07-24</link>
            <guid>http://easyai.fyi/article/follow-builders-ai-summary-2026-07-24</guid>
            <pubDate>Fri, 24 Jul 2026 00:00:00 GMT</pubDate>
            <description><![CDATA[本轮 follow-builders 新增 12 个 X builders、27 条 tweet、1 篇 blog 和 1 期 podcast，已按输出全量沉淀。今天主线是 Agent 正在进入语音、开发、团队协作和自主业务流程，同时 inference 速度、芯片和开放模型继续成为底层竞争焦点。]]></description>
            <content:encoded><![CDATA[<div id="notion-article" class="mx-auto overflow-hidden "><main class="notion light-mode notion-page notion-block-3a75eac3813c81909140d9cc795d0b35"><div class="notion-viewport"></div><div class="notion-collection-page-properties"></div><h3 class="notion-h notion-h2 notion-h-indent-0 notion-block-e372698f9fc44c7b97f8c313135a603f" data-id="e372698f9fc44c7b97f8c313135a603f"><span><div id="e372698f9fc44c7b97f8c313135a603f" class="notion-header-anchor"></div><a class="notion-hash-link" href="#e372698f9fc44c7b97f8c313135a603f" title="今日主线"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">今日主线</span></span></h3><div class="notion-text notion-block-e7b3312493784ab98a44ed9eca5e5166">今天这批内容的主线不是单个模型发布，而是 AI 正在进入更具体的工作场景：ChatGPT 和 Claude 都在把 voice 变成更自然的操作入口，Replit、Claude Code、Vercel 这类工具继续把 agent 拉进真实工程流程；另一边，Cerebras 的 Andrew Feldman 把焦点放到 inference 速度、memory、制造和电力上，提醒大家 AI 竞争不是只看模型参数，也看谁能把响应做得更快、更便宜、更稳定。</div><h3 class="notion-h notion-h2 notion-h-indent-0 notion-block-b9591c8dcab1470295273c9455a8ef8d" data-id="b9591c8dcab1470295273c9455a8ef8d"><span><div id="b9591c8dcab1470295273c9455a8ef8d" class="notion-header-anchor"></div><a class="notion-hash-link" href="#b9591c8dcab1470295273c9455a8ef8d" title="重点解读"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">重点解读</span></span></h3><div class="notion-text notion-block-e6a0133d548741d4924abf4862bf1411">最值得看的是 voice 和 agent 工作流开始合流。OpenAI 侧的 Thibault Sottiaux 在推 ChatGPT desktop app 里的离键盘工作方式，Peter Yang 也在试 ChatGPT Voice，并提出下一步可能是多条 voice thread 并行协作。Claude 这边则把 voice mode 接到更强模型和已连接工具上，支持更多语言。放在一起看，AI 产品正在从“打字问答”转向“边说边派活”，这会明显改变日常工作入口。</div><div class="notion-text notion-block-e1e3067b96ba412885b47f1edc1d7785">工程工具这条线也在变厚。Swyx 提到自己在 dogfood 一个 agentic GitHub clone，还强调 Poolside AI 不只发布模型，也公开完整 eval dataset。Amjad Masad 讲 Replit 用户把 agency 业务拆成 agent loop，并通过 MCP 进一步自动化。Claude Blog 则发布 Claude Code artifacts，让一次 coding session 的调查、PR 解释、dashboard 或 release checklist 变成可分享、会更新的页面。这里的共同点是：Agent 不再只是生成代码，而是在把工作过程本身产品化。</div><div class="notion-text notion-block-de586a7d42e44459ae2638c527f4b705">安全和权限问题也开始变得很具体。Madhu Guru 提到一个很现实的问题：过去 IAM 是给有限员工设计的，但现在一个员工可以启动上百个 agent，agent 还能再启动子 agent。权限继承、生命周期、审计边界都会变复杂。Aaron Levie 的判断也接近这个方向：AI 更像已有专业能力的放大器，专家能用它做更多高质量工作，没有判断力的人反而更容易产出低质量内容。</div><div class="notion-text notion-block-575db8907ba746f7a0153b16eae3c817">底层基建这边，Matt Turck 和 Cerebras CEO Andrew Feldman 的对话很适合和今天的 agent 讨论放在一起。Feldman 的核心判断是，AI 从新奇玩具变成生产力工具后，速度会直接变成体验和价值。他把关键指标说成 “tokens per second per user”，并强调 agent 多轮调用会放大等待时间。访谈还提到 HBM、CoWoS、TSMC 先进制程、电力和中国开源模型这些更底层的约束，说明 inference 不是简单买 GPU 就完事。</div><h3 class="notion-h notion-h2 notion-h-indent-0 notion-block-c49578b2998041f0a539aa789d4b8fe6" data-id="c49578b2998041f0a539aa789d4b8fe6"><span><div id="c49578b2998041f0a539aa789d4b8fe6" class="notion-header-anchor"></div><a class="notion-hash-link" href="#c49578b2998041f0a539aa789d4b8fe6" title="X Builders 全量记录"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">X Builders 全量记录</span></span></h3><h4 class="notion-h notion-h3 notion-h-indent-1 notion-block-df5470ffd5ac4082a6f3c1621583126f" data-id="df5470ffd5ac4082a6f3c1621583126f"><span><div id="df5470ffd5ac4082a6f3c1621583126f" class="notion-header-anchor"></div><a class="notion-hash-link" href="#df5470ffd5ac4082a6f3c1621583126f" title="1. Swyx @swyx"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">1. Swyx @swyx</span></span></h4><div class="notion-text notion-block-ba9bf5fa43b04fe09dae592d52532445">Swyx 今天两条都在工程主线里。一条是他过去一个月在 dogfood 一个 agentic GitHub clone，并提到已经做到内置 CI/CD，说明代码托管、自动化部署和 agent 工作流正在被揉到一起；另一条是夸 Poolside AI 的开放程度，重点不是只说模型成绩，而是他们把 eval dataset 也公开出来，让外界能自己看有没有 reward hacking。这类透明度对开发者很重要，因为模型竞争正在从“榜单分数”走向“评测过程是否可信”。</div><figure class="notion-asset-wrapper notion-asset-wrapper-embed notion-block-474f5eaacb9c4ba38a45ca33998bf8e3"><div style="position:relative;display:flex;justify-content:center;align-self:center;width:100%;max-width:100%;flex-direction:column;height:320px"><iframe class="notion-asset-object-fit" src="https://platform.twitter.com/embed/Tweet.html?id=2080500752183960017" title="iframe embed" frameBorder="0" allowfullscreen="" loading="lazy" scrolling="auto"></iframe></div></figure><figure class="notion-asset-wrapper notion-asset-wrapper-embed notion-block-de693b43ac0e4fa69ef96bc492846d8b"><div style="position:relative;display:flex;justify-content:center;align-self:center;width:100%;max-width:100%;flex-direction:column;height:320px"><iframe class="notion-asset-object-fit" src="https://platform.twitter.com/embed/Tweet.html?id=2080387171723137440" title="iframe embed" frameBorder="0" allowfullscreen="" loading="lazy" scrolling="auto"></iframe></div></figure><h4 class="notion-h notion-h3 notion-h-indent-1 notion-block-472cff93b57842b6b25a1be114a687cb" data-id="472cff93b57842b6b25a1be114a687cb"><span><div id="472cff93b57842b6b25a1be114a687cb" class="notion-header-anchor"></div><a class="notion-hash-link" href="#472cff93b57842b6b25a1be114a687cb" title="2. Thibault Sottiaux @thsottiaux"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">2. Thibault Sottiaux @thsottiaux</span></span></h4><div class="notion-text notion-block-1cc515729e5b4b24b4e8ca674d5abf32">Thibault 的三条围绕 ChatGPT Work 和桌面端体验。一条像是在试探 ChatGPT Work 是否该改名成 ChatGPT Vibe，另一条偏招聘和愿景表达，最有信息量的是他把新的离键盘工作方式和 Jarvis、Samantha、TARS 这类科幻助手类比，并说已经可以在 ChatGPT desktop app 里尝试。它传递的信号很直接：OpenAI 想把 ChatGPT 从聊天窗口推向更持续、更自然的工作入口。</div><figure class="notion-asset-wrapper notion-asset-wrapper-embed notion-block-b80b882b6bef4772b3012f3dc9029b33"><div style="position:relative;display:flex;justify-content:center;align-self:center;width:100%;max-width:100%;flex-direction:column;height:320px"><iframe class="notion-asset-object-fit" src="https://platform.twitter.com/embed/Tweet.html?id=2080543574211666029" title="iframe embed" frameBorder="0" allowfullscreen="" loading="lazy" scrolling="auto"></iframe></div></figure><figure class="notion-asset-wrapper notion-asset-wrapper-embed notion-block-769791113bd04eb5a77c12a9550d6618"><div style="position:relative;display:flex;justify-content:center;align-self:center;width:100%;max-width:100%;flex-direction:column;height:320px"><iframe class="notion-asset-object-fit" src="https://platform.twitter.com/embed/Tweet.html?id=2080537149204758689" title="iframe embed" frameBorder="0" allowfullscreen="" loading="lazy" scrolling="auto"></iframe></div></figure><figure class="notion-asset-wrapper notion-asset-wrapper-embed notion-block-8ee9da1c015340e1b58b6c5011569097"><div style="position:relative;display:flex;justify-content:center;align-self:center;width:100%;max-width:100%;flex-direction:column;height:320px"><iframe class="notion-asset-object-fit" src="https://platform.twitter.com/embed/Tweet.html?id=2080408012515340394" title="iframe embed" frameBorder="0" allowfullscreen="" loading="lazy" scrolling="auto"></iframe></div></figure><h4 class="notion-h notion-h3 notion-h-indent-1 notion-block-ea5cb1444a724a2da70c816823589881" data-id="ea5cb1444a724a2da70c816823589881"><span><div id="ea5cb1444a724a2da70c816823589881" class="notion-header-anchor"></div><a class="notion-hash-link" href="#ea5cb1444a724a2da70c816823589881" title="3. Peter Yang @petergyang"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">3. Peter Yang @petergyang</span></span></h4><div class="notion-text notion-block-7e4f55ae32cc43dd8363e796505808be">Peter Yang 也在围绕 ChatGPT Voice 反馈产品体验。他觉得下一步会是同时开多个 ChatGPT Voice threads，让一组“人”在你旁边协作、互相对话；他也提到多线程完成时需要提醒，以及中文发音还不够好。这些反馈很产品化，重点不是 voice 好不好玩，而是 voice 一旦能承载多任务和协作提醒，就可能从输入方式变成 agent 调度界面。</div><figure class="notion-asset-wrapper notion-asset-wrapper-embed notion-block-e147562678974ee2ad42da4bc2e8ab06"><div style="position:relative;display:flex;justify-content:center;align-self:center;width:100%;max-width:100%;flex-direction:column;height:320px"><iframe class="notion-asset-object-fit" src="https://platform.twitter.com/embed/Tweet.html?id=2080508139091427741" title="iframe embed" frameBorder="0" allowfullscreen="" loading="lazy" scrolling="auto"></iframe></div></figure><figure class="notion-asset-wrapper notion-asset-wrapper-embed notion-block-00d9880d5430480fb63f0d12ec26849b"><div style="position:relative;display:flex;justify-content:center;align-self:center;width:100%;max-width:100%;flex-direction:column;height:320px"><iframe class="notion-asset-object-fit" src="https://platform.twitter.com/embed/Tweet.html?id=2080505964936241226" title="iframe embed" frameBorder="0" allowfullscreen="" loading="lazy" scrolling="auto"></iframe></div></figure><figure class="notion-asset-wrapper notion-asset-wrapper-embed notion-block-c55ae70a66bd4387a06a4f4c145c2a6f"><div style="position:relative;display:flex;justify-content:center;align-self:center;width:100%;max-width:100%;flex-direction:column;height:320px"><iframe class="notion-asset-object-fit" src="https://platform.twitter.com/embed/Tweet.html?id=2080505108216111303" title="iframe embed" frameBorder="0" allowfullscreen="" loading="lazy" scrolling="auto"></iframe></div></figure><h4 class="notion-h notion-h3 notion-h-indent-1 notion-block-b5cec6f15fc8412ebf04a1931a386007" data-id="b5cec6f15fc8412ebf04a1931a386007"><span><div id="b5cec6f15fc8412ebf04a1931a386007" class="notion-header-anchor"></div><a class="notion-hash-link" href="#b5cec6f15fc8412ebf04a1931a386007" title="4. Madhu Guru @realmadhuguru"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">4. Madhu Guru @realmadhuguru</span></span></h4><div class="notion-text notion-block-42c7b9b7ec4b458cbd3c79b121d6e7f7">Madhu Guru 一条讲 AI 模型的 jagged frontier，另一条更关键：他和一家上市公司安全负责人聊 GPT Sol incident 后的安全问题，焦点落在“无限 agent”时代的身份和权限管理。传统系统假设员工数量有限，每个人有身份、角色、权限和生命周期；但如果一个员工能启动上百个 agent，子 agent 又继续派生，权限继承、审计和生命周期就会成为企业 AI 的硬问题。这是今天最值得企业团队记下的一条。</div><figure class="notion-asset-wrapper notion-asset-wrapper-embed notion-block-57cbeaf7c1d0446c8ad612eeca84f677"><div style="position:relative;display:flex;justify-content:center;align-self:center;width:100%;max-width:100%;flex-direction:column;height:320px"><iframe class="notion-asset-object-fit" src="https://platform.twitter.com/embed/Tweet.html?id=2080460579966501257" title="iframe embed" frameBorder="0" allowfullscreen="" loading="lazy" scrolling="auto"></iframe></div></figure><figure class="notion-asset-wrapper notion-asset-wrapper-embed notion-block-30e9bbd4625140b7812f3dc89e301401"><div style="position:relative;display:flex;justify-content:center;align-self:center;width:100%;max-width:100%;flex-direction:column;height:320px"><iframe class="notion-asset-object-fit" src="https://platform.twitter.com/embed/Tweet.html?id=2080315474093760714" title="iframe embed" frameBorder="0" allowfullscreen="" loading="lazy" scrolling="auto"></iframe></div></figure><h4 class="notion-h notion-h3 notion-h-indent-1 notion-block-c66e0b974fd147c7b6d795f90ba91cdb" data-id="c66e0b974fd147c7b6d795f90ba91cdb"><span><div id="c66e0b974fd147c7b6d795f90ba91cdb" class="notion-header-anchor"></div><a class="notion-hash-link" href="#c66e0b974fd147c7b6d795f90ba91cdb" title="5. Amjad Masad @amasad"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">5. Amjad Masad @amasad</span></span></h4><div class="notion-text notion-block-269b11fb69b4489da07393e91cde9e03">Amjad Masad 的三条都和 Replit 的产品和 agent 化有关。他提到 autoscale deployments 成本下降 80%，又说自己的 chess autoresearch agent 像是学会了现代 LLM fine-tuning；更有意思的是一个用户用 Replit 先打破传统 agency 模式，随后又想把整个 agency 自动化，于是通过 MCP 建出 autonomous agency。这里的信号很清楚：agent loop 不只发生在写代码环节，也会扩展到接单、交付、运营这一整套业务流程。</div><figure class="notion-asset-wrapper notion-asset-wrapper-embed notion-block-ab9cf99afa834424b8dd492b2fa0e7be"><div style="position:relative;display:flex;justify-content:center;align-self:center;width:100%;max-width:100%;flex-direction:column;height:320px"><iframe class="notion-asset-object-fit" src="https://platform.twitter.com/embed/Tweet.html?id=2080513361301925957" title="iframe embed" frameBorder="0" allowfullscreen="" loading="lazy" scrolling="auto"></iframe></div></figure><figure class="notion-asset-wrapper notion-asset-wrapper-embed notion-block-6b97b51edfbb4101abb3ef288b6c3c09"><div style="position:relative;display:flex;justify-content:center;align-self:center;width:100%;max-width:100%;flex-direction:column;height:320px"><iframe class="notion-asset-object-fit" src="https://platform.twitter.com/embed/Tweet.html?id=2080512523389005894" title="iframe embed" frameBorder="0" allowfullscreen="" loading="lazy" scrolling="auto"></iframe></div></figure><figure class="notion-asset-wrapper notion-asset-wrapper-embed notion-block-8a52bebcbc814fa4aaa2d9c5b805d841"><div style="position:relative;display:flex;justify-content:center;align-self:center;width:100%;max-width:100%;flex-direction:column;height:320px"><iframe class="notion-asset-object-fit" src="https://platform.twitter.com/embed/Tweet.html?id=2080371567221944657" title="iframe embed" frameBorder="0" allowfullscreen="" loading="lazy" scrolling="auto"></iframe></div></figure><h4 class="notion-h notion-h3 notion-h-indent-1 notion-block-5b64d67f44ef4520bf63ceaefdf7ade7" data-id="5b64d67f44ef4520bf63ceaefdf7ade7"><span><div id="5b64d67f44ef4520bf63ceaefdf7ade7" class="notion-header-anchor"></div><a class="notion-hash-link" href="#5b64d67f44ef4520bf63ceaefdf7ade7" title="6. Guillermo Rauch @rauchg"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">6. Guillermo Rauch @rauchg</span></span></h4><div class="notion-text notion-block-1a42bc5fbcba4bcd833073a3c7b63de1">Guillermo Rauch 今天主要围绕 Vercel 的开发者基建。一条说 Python code 在 Vercel 上启动速度提升到 2 倍，另一条说 AI Gateway 的产品速度很快。和今天的 Cerebras 访谈放在一起看，这也是同一条线：AI 应用的体验最终会回到延迟、冷启动、token 成本、gateway 和部署效率这些基础设施细节上。</div><figure class="notion-asset-wrapper notion-asset-wrapper-embed notion-block-825ed767128a45728a5865b4841815ea"><div style="position:relative;display:flex;justify-content:center;align-self:center;width:100%;max-width:100%;flex-direction:column;height:320px"><iframe class="notion-asset-object-fit" src="https://platform.twitter.com/embed/Tweet.html?id=2080454509508387251" title="iframe embed" frameBorder="0" allowfullscreen="" loading="lazy" scrolling="auto"></iframe></div></figure><figure class="notion-asset-wrapper notion-asset-wrapper-embed notion-block-ff94298b9b874042b0ad60f3ab0ad131"><div style="position:relative;display:flex;justify-content:center;align-self:center;width:100%;max-width:100%;flex-direction:column;height:320px"><iframe class="notion-asset-object-fit" src="https://platform.twitter.com/embed/Tweet.html?id=2080344136625049690" title="iframe embed" frameBorder="0" allowfullscreen="" loading="lazy" scrolling="auto"></iframe></div></figure><h4 class="notion-h notion-h3 notion-h-indent-1 notion-block-963150e4baa54377b9a5e60ba35f625a" data-id="963150e4baa54377b9a5e60ba35f625a"><span><div id="963150e4baa54377b9a5e60ba35f625a" class="notion-header-anchor"></div><a class="notion-hash-link" href="#963150e4baa54377b9a5e60ba35f625a" title="7. Aaron Levie @levie"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">7. Aaron Levie @levie</span></span></h4><div class="notion-text notion-block-3ceefcad61424c0b9199cb999257ca1c">Aaron Levie 把 AI 看成已有专业能力的放大器，而不是替代判断力的魔法。他的判断是，真正会产生经济价值的是专家用 AI 做更多、更高质量的工作，因为他们知道怎么把 agent 拉回正确方向，并把结果接进真实工作里；没有判断力、也不想建立判断力的人，更容易制造 slop。这是对今天 agent 热潮的一个冷静提醒：工具越强，人的判断越值钱。</div><figure class="notion-asset-wrapper notion-asset-wrapper-embed notion-block-6db85f2f3cb545b5ae313229e159628f"><div style="position:relative;display:flex;justify-content:center;align-self:center;width:100%;max-width:100%;flex-direction:column;height:320px"><iframe class="notion-asset-object-fit" src="https://platform.twitter.com/embed/Tweet.html?id=2080471989060559336" title="iframe embed" frameBorder="0" allowfullscreen="" loading="lazy" scrolling="auto"></iframe></div></figure><h4 class="notion-h notion-h3 notion-h-indent-1 notion-block-fd38c7e9f858486ba792701cb4ee930d" data-id="fd38c7e9f858486ba792701cb4ee930d"><span><div id="fd38c7e9f858486ba792701cb4ee930d" class="notion-header-anchor"></div><a class="notion-hash-link" href="#fd38c7e9f858486ba792701cb4ee930d" title="8. Garry Tan @garrytan"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">8. Garry Tan @garrytan</span></span></h4><div class="notion-text notion-block-424e71e829f34145b1e51e48a0da2fd0">Garry Tan 今天有两条关于旧金山住房和 CEQA 改革，属于非 AI 主线，但按全量要求保留。另一条与 AI 有关，他强调 open weight models 非常重要。它和 Swyx 提到 Poolside 开放 eval dataset 可以放在一起看：开放权重、开放评测、开放复现，仍然是 AI 生态里对抗黑箱化的重要力量。</div><figure class="notion-asset-wrapper notion-asset-wrapper-embed notion-block-bcd7360913ec4c569b8b326a39c4dfb7"><div style="position:relative;display:flex;justify-content:center;align-self:center;width:100%;max-width:100%;flex-direction:column;height:320px"><iframe class="notion-asset-object-fit" src="https://platform.twitter.com/embed/Tweet.html?id=2080443154730553402" title="iframe embed" frameBorder="0" allowfullscreen="" loading="lazy" scrolling="auto"></iframe></div></figure><figure class="notion-asset-wrapper notion-asset-wrapper-embed notion-block-96b6d1e2fdb24d54af268ff824a7db65"><div style="position:relative;display:flex;justify-content:center;align-self:center;width:100%;max-width:100%;flex-direction:column;height:320px"><iframe class="notion-asset-object-fit" src="https://platform.twitter.com/embed/Tweet.html?id=2080364752778527195" title="iframe embed" frameBorder="0" allowfullscreen="" loading="lazy" scrolling="auto"></iframe></div></figure><figure class="notion-asset-wrapper notion-asset-wrapper-embed notion-block-ee50e5c8e04e456ca1c1e1384c1555d6"><div style="position:relative;display:flex;justify-content:center;align-self:center;width:100%;max-width:100%;flex-direction:column;height:320px"><iframe class="notion-asset-object-fit" src="https://platform.twitter.com/embed/Tweet.html?id=2080345524620914897" title="iframe embed" frameBorder="0" allowfullscreen="" loading="lazy" scrolling="auto"></iframe></div></figure><h4 class="notion-h notion-h3 notion-h-indent-1 notion-block-46da0239bea9475b96d33edc72d5cce5" data-id="46da0239bea9475b96d33edc72d5cce5"><span><div id="46da0239bea9475b96d33edc72d5cce5" class="notion-header-anchor"></div><a class="notion-hash-link" href="#46da0239bea9475b96d33edc72d5cce5" title="9. Matt Turck @mattturck"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">9. Matt Turck @mattturck</span></span></h4><div class="notion-text notion-block-5f40344bb81f415cb75f7f0f90a8cd42">Matt Turck 今天一条是关于 VC 和高算力 AI 公司的调侃，另外两条都指向他和 Cerebras CEO Andrew Feldman 的 fast inference、AI chips、下一代 compute bottleneck 对话。他列出的主题很完整：tokens per second per user、GPU/TPU/Trainium/ASIC、OpenAI 与专用芯片、中国与主权 AI 基建、HBM、CoWoS、3nm、以及 agents 为什么会带来 CPU 需求。今天 podcast 的主线也基本围绕这些展开。</div><figure class="notion-asset-wrapper notion-asset-wrapper-embed notion-block-ab06aed6870340d5b416c3f364ccc136"><div style="position:relative;display:flex;justify-content:center;align-self:center;width:100%;max-width:100%;flex-direction:column;height:320px"><iframe class="notion-asset-object-fit" src="https://platform.twitter.com/embed/Tweet.html?id=2080451010439352711" title="iframe embed" frameBorder="0" allowfullscreen="" loading="lazy" scrolling="auto"></iframe></div></figure><figure class="notion-asset-wrapper notion-asset-wrapper-embed notion-block-4fef27ed941d4654b485658e82804a2e"><div style="position:relative;display:flex;justify-content:center;align-self:center;width:100%;max-width:100%;flex-direction:column;height:320px"><iframe class="notion-asset-object-fit" src="https://platform.twitter.com/embed/Tweet.html?id=2080333711640285549" title="iframe embed" frameBorder="0" allowfullscreen="" loading="lazy" scrolling="auto"></iframe></div></figure><figure class="notion-asset-wrapper notion-asset-wrapper-embed notion-block-256c9843141b4f7c8c9f49bfbc7f7c84"><div style="position:relative;display:flex;justify-content:center;align-self:center;width:100%;max-width:100%;flex-direction:column;height:320px"><iframe class="notion-asset-object-fit" src="https://platform.twitter.com/embed/Tweet.html?id=2080333707483725876" title="iframe embed" frameBorder="0" allowfullscreen="" loading="lazy" scrolling="auto"></iframe></div></figure><h4 class="notion-h notion-h3 notion-h-indent-1 notion-block-db43c4ff333a49a6af7f8cebe5d5341c" data-id="db43c4ff333a49a6af7f8cebe5d5341c"><span><div id="db43c4ff333a49a6af7f8cebe5d5341c" class="notion-header-anchor"></div><a class="notion-hash-link" href="#db43c4ff333a49a6af7f8cebe5d5341c" title="10. Nikunj Kothari @nikunj"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">10. Nikunj Kothari @nikunj</span></span></h4><div class="notion-text notion-block-5af4a75f6e5b446db3e51a11ffc95740">Nikunj Kothari 这条是对科技行业词汇膨胀的吐槽。他列了 “neo-something”、full stack、fellows、labs、partner、forward deployed、RL 等越来越被滥用的标签。它不是 AI 产品更新，但和当下 AI 行业的命名焦虑有关：当所有东西都被包装成新范式时，信号反而会变弱。</div><figure class="notion-asset-wrapper notion-asset-wrapper-embed notion-block-66ba6c70431a45a6a70646e285a9e71b"><div style="position:relative;display:flex;justify-content:center;align-self:center;width:100%;max-width:100%;flex-direction:column;height:320px"><iframe class="notion-asset-object-fit" src="https://platform.twitter.com/embed/Tweet.html?id=2080293627784212933" title="iframe embed" frameBorder="0" allowfullscreen="" loading="lazy" scrolling="auto"></iframe></div></figure><h4 class="notion-h notion-h3 notion-h-indent-1 notion-block-72073ec0970d41c4ae4b31d56d2bc9b4" data-id="72073ec0970d41c4ae4b31d56d2bc9b4"><span><div id="72073ec0970d41c4ae4b31d56d2bc9b4" class="notion-header-anchor"></div><a class="notion-hash-link" href="#72073ec0970d41c4ae4b31d56d2bc9b4" title="11. Peter Steinberger @steipete"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">11. Peter Steinberger @steipete</span></span></h4><div class="notion-text notion-block-b941a2b2873f483d98c8448503450be0">Peter Steinberger 这条是很实用的工程反馈：他们也遇到相关问题，并加了直接使用 claude cli 的 code paths。他没有展开太多背景，但对 agent/harness 开发者来说，这是一个典型现实：当外层系统不好控制时，团队会绕到更直接的 CLI 路径上，优先保证任务能跑通。</div><figure class="notion-asset-wrapper notion-asset-wrapper-embed notion-block-150a3226a5664d31bb1204a9f1fa8b30"><div style="position:relative;display:flex;justify-content:center;align-self:center;width:100%;max-width:100%;flex-direction:column;height:320px"><iframe class="notion-asset-object-fit" src="https://platform.twitter.com/embed/Tweet.html?id=2080318789980201224" title="iframe embed" frameBorder="0" allowfullscreen="" loading="lazy" scrolling="auto"></iframe></div></figure><h4 class="notion-h notion-h3 notion-h-indent-1 notion-block-1c5c374d195c40bab468f247630225ce" data-id="1c5c374d195c40bab468f247630225ce"><span><div id="1c5c374d195c40bab468f247630225ce" class="notion-header-anchor"></div><a class="notion-hash-link" href="#1c5c374d195c40bab468f247630225ce" title="12. Claude @claudeai"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">12. Claude @claudeai</span></span></h4><div class="notion-text notion-block-8af0decae5ad4bcebdd4eb8cf88506d6">Claude 官方三条都围绕 voice mode 更新。新的 voice conversations 可以使用更多 chat 里的模型，包括 Claude Opus 和 Sonnet，也能在对话中调用用户已连接的工具，比如 email 和 calendar；同时 voice mode 支持更多语言，并在 mobile、desktop、web 上以 public beta 推出。这说明 Claude 也在把语音从“对话能力”推向“带工具的工作入口”。</div><figure class="notion-asset-wrapper notion-asset-wrapper-embed notion-block-649c5780a2c541479de161ea3b695b03"><div style="position:relative;display:flex;justify-content:center;align-self:center;width:100%;max-width:100%;flex-direction:column;height:320px"><iframe class="notion-asset-object-fit" src="https://platform.twitter.com/embed/Tweet.html?id=2080376099268169943" title="iframe embed" frameBorder="0" allowfullscreen="" loading="lazy" scrolling="auto"></iframe></div></figure><figure class="notion-asset-wrapper notion-asset-wrapper-embed notion-block-8e4a9999bc2b4cb9b8da5fcfa36f3366"><div style="position:relative;display:flex;justify-content:center;align-self:center;width:100%;max-width:100%;flex-direction:column;height:320px"><iframe class="notion-asset-object-fit" src="https://platform.twitter.com/embed/Tweet.html?id=2080376096873177300" title="iframe embed" frameBorder="0" allowfullscreen="" loading="lazy" scrolling="auto"></iframe></div></figure><figure class="notion-asset-wrapper notion-asset-wrapper-embed notion-block-4ff78be84ff647d3865ff6763446bed1"><div style="position:relative;display:flex;justify-content:center;align-self:center;width:100%;max-width:100%;flex-direction:column;height:320px"><iframe class="notion-asset-object-fit" src="https://platform.twitter.com/embed/Tweet.html?id=2080376094939603366" title="iframe embed" frameBorder="0" allowfullscreen="" loading="lazy" scrolling="auto"></iframe></div></figure><h3 class="notion-h notion-h2 notion-h-indent-0 notion-block-c17257a196d14fa8be315d7301fd2baa" data-id="c17257a196d14fa8be315d7301fd2baa"><span><div id="c17257a196d14fa8be315d7301fd2baa" class="notion-header-anchor"></div><a class="notion-hash-link" href="#c17257a196d14fa8be315d7301fd2baa" title="Blog 全量记录"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">Blog 全量记录</span></span></h3><h4 class="notion-h notion-h3 notion-h-indent-1 notion-block-8b15cf008f9a4516b65c12efd4e07f53" data-id="8b15cf008f9a4516b65c12efd4e07f53"><span><div id="8b15cf008f9a4516b65c12efd4e07f53" class="notion-header-anchor"></div><a class="notion-hash-link" href="#8b15cf008f9a4516b65c12efd4e07f53" title="1. Claude Blog：Claude Code now supports artifacts"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">1. Claude Blog：Claude Code now supports artifacts</span></span></h4><div class="notion-text notion-block-d5551a04c31e4c0da66db69a9d59bf11">Claude Code artifacts 的核心是把一次 Claude Code session 的进展变成可打开、可分享、会更新的 live page，比如 PR walkthrough、system explainer、dashboard、release checklist、incident timeline。它利用 session 里的 codebase、connectors 和对话上下文生成页面，不需要团队先搭数据源或站点；后续更新会发布到同一个链接，并保留版本历史。权限上，artifact 默认只对作者私有，分享后也只给组织内认证成员访问，管理员可以做组织级开关、角色范围、保留策略和合规 API 可见性。Beta 面向 Claude Team 和 Enterprise 组织，支持 Claude Code CLI 和 desktop app。原文：<a target="_blank" rel="noopener noreferrer" class="notion-link" href="https://claude.com/blog/artifacts-in-claude-code">Claude Code now supports artifacts</a></div><h3 class="notion-h notion-h2 notion-h-indent-0 notion-block-a356dc1757514d7d8caad5c9ae9e5d81" data-id="a356dc1757514d7d8caad5c9ae9e5d81"><span><div id="a356dc1757514d7d8caad5c9ae9e5d81" class="notion-header-anchor"></div><a class="notion-hash-link" href="#a356dc1757514d7d8caad5c9ae9e5d81" title="Podcast 全量记录"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">Podcast 全量记录</span></span></h3><h4 class="notion-h notion-h3 notion-h-indent-1 notion-block-95d613924f9c44e0a2db56c96d9a9f59" data-id="95d613924f9c44e0a2db56c96d9a9f59"><span><div id="95d613924f9c44e0a2db56c96d9a9f59" class="notion-header-anchor"></div><a class="notion-hash-link" href="#95d613924f9c44e0a2db56c96d9a9f59" title="1. The MAD Podcast with Matt Turck：The Biggest Chip Ever Built — Why OpenAI Runs On It | Cerebras CEO Andrew Feldman"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">1. The MAD Podcast with Matt Turck：The Biggest Chip Ever Built — Why OpenAI Runs On It | Cerebras CEO Andrew Feldman</span></span></h4><div class="notion-text notion-block-9b3128da812e4d92b03e81dbb46291ad">The MAD Podcast 这期最重要的 takeaway 是：AI 进入真实生产力场景后，速度会变成核心竞争变量。Cerebras CEO Andrew Feldman 的说法很直接，AI 是用 training 做出来的，但真正使用时靠 inference；当 agent 有多轮调用、验证和 guardrails 时，等待会被放大，所以关键指标不是抽象算力，而是每个用户每秒能拿到多少 token。访谈还把芯片格局拆开讲：GPU、TPU、Trainium、ASIC 都是在为不同 AI workload 做选择；Cerebras 的卖点是超大 wafer-scale chip，号称比 GPU 大 58 倍，并试图绕开 HBM、CoWoS 和 3nm 等供给瓶颈。更值得留意的是他对 agent 的判断：大模型可以先产出答案，小模型再像 fact checker 一样做核查；这会继续增加 inference 和验证成本。来源：<a target="_blank" rel="noopener noreferrer" class="notion-link" href="https://www.youtube.com/@DataDrivenNYC/videos">YouTube 原始链接</a></div><h3 class="notion-h notion-h2 notion-h-indent-0 notion-block-4fab87b72aa94b2fab4c5b3cd7ae320f" data-id="4fab87b72aa94b2fab4c5b3cd7ae320f"><span><div id="4fab87b72aa94b2fab4c5b3cd7ae320f" class="notion-header-anchor"></div><a class="notion-hash-link" href="#4fab87b72aa94b2fab4c5b3cd7ae320f" title="今日沉淀结论"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">今日沉淀结论</span></span></h3><div class="notion-text notion-block-92f52c539f7c4981802c417c54645790">今天所有内容放在一起看，AI 正在同时向上和向下推进：向上是 voice、agent、MCP、artifacts 把工作流变得更自动；向下是 inference、芯片、memory、电力和部署成本决定这些体验能不能真的跑起来。下一阶段的竞争，大概率不是“谁会聊天”，而是谁能把 agent 放进真实业务里，并用足够快、足够便宜、足够可控的基础设施撑住它。</div><div class="notion-text notion-block-205b3ca5423c4c8e9d73af3316f1c75f">Generated through the Follow Builders skill: <a target="_blank" rel="noopener noreferrer" class="notion-link" href="https://github.com/zarazhangrui/follow-builders">https://github.com/zarazhangrui/follow-builders</a></div></main></div>]]></content:encoded>
        </item>
    </channel>
</rss>