<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/"><channel><title>Evals on AI Digest</title><link>https://aidigest.kikihuang.net/tags/evals/</link><description>Recent content in Evals on AI Digest</description><generator>Hugo</generator><language>zh-cn</language><lastBuildDate>Wed, 08 Jul 2026 00:02:00 +0800</lastBuildDate><atom:link href="https://aidigest.kikihuang.net/tags/evals/index.xml" rel="self" type="application/rss+xml"/><item><title>AI Builders 日报 — 2026年7月8日</title><link>https://aidigest.kikihuang.net/posts/2026-07-08-daily/</link><pubDate>Wed, 08 Jul 2026 00:02:00 +0800</pubDate><guid>https://aidigest.kikihuang.net/posts/2026-07-08-daily/</guid><description>follow-builders 的 X / Podcast 源本期正常抓取（16 位 builder、34 条推文、1 集播客、1 篇 blog）。今日主线极其集中——Anthropic 与 @claudeai 官方同步放出《Claude Code 起源史》，Boris Cherny 与 Cat Wu 讲述它如何从『安全研究内部工具』长成产品；同日 Anthropic 发布 J-space 可解释性论文，Swyx 划重点：他们能对模型推理做『脑外科手术』式干预，且模型能『察觉自己被干预了』，逼近 eval awareness。另一条支线是『agent 开始自我进化 / 自我评测』：Replit 宣称已闭环让 agent self-improving，Vercel 的 eve 用 &lt;code&gt;eve eval&lt;/code&gt; 给自己做进化评测。Fable 5 今晚 23:59 PT 下线，Peter Yang 给出最后 5 个值得一试的高价值 prompt。</description></item><item><title>AI Builders 日报 — 2026年7月7日</title><link>https://aidigest.kikihuang.net/posts/2026-07-07-daily/</link><pubDate>Tue, 07 Jul 2026 00:00:00 +0800</pubDate><guid>https://aidigest.kikihuang.net/posts/2026-07-07-daily/</guid><description>follow-builders 的 X / Podcast 源本期恢复正常抓取（14 位 builder、25 条推文、1 集播客）。主线是「Agent 从炫技走向真实工作流」：Cat Wu 用 Claude Code 自动 sourcing 候选人并邮寄清单，Nan Yu 直指『开 10 个 Claude 标签页只是表演』；播客侧 No Priors 请到 OpenAI 的 Noam Brown，系统讲清 test-time compute 如何改写 benchmark、safety 评测与 RSI 的节奏——他认为不会有『一夜之间的智能爆炸』，因为最强能力被算力与时间卡住。</description></item></channel></rss>