我用七步法继续调试 Grok Bot,三只 Bot 终于组成了一条内容生产线

上一篇写 Grok Bot 时,我在结尾说,暂时不继续创建 Content Scout 或 Chief of Staff,先让 Theatre Scout、AI Signal Radar 和 Brand 5.0 Lab Researcher 跑上 2 到 4 周。

结果很快,我又打开了 Grok Bot。

倒不是因为我突然想凑一支浩浩荡荡的 Bot 队伍,而是第一批 Bot 跑起来之后,一个缺口变得很明显:我已经有了找演出的、追 AI 信号的、做 Brand 5.0 Lab 研究的 Bot,但这些信息最后还是全部堆回到我这里。

我要自己看视频、核对原始材料、判断哪些能写、哪些只是一家之言,再把 X 上的信号、视频知识和自己的观点拼成一篇文章。“科爷的数字生命”这个账号,真正需要的不是又多一个会聊天的数字分身,而是有人能把这些环节接起来。

于是,我又做了三只 Bot:

  • • X Intelligence Scout,从 X 发现与我真正相关的高信号。

  • • Video Knowledge Editor,把视频处理成有来源、可追溯的知识资产。

  • • Idea to Article Editor,把不同来源的材料组织成有明确 thesis 的文章。

它们最后形成了一条很清楚的链条:

Signal / Research
      ↓
Source Processing
      ↓
Editorial Synthesis
      ↓
Final Article Draft
01-framework-content-pipeline.png

看上去只是三个岗位。但我后来花得最多的时间,仍然不是创建,而是调试。

第一次做对,只能证明它会工作

上一篇里,我把调试 Grok Bot 总结成七步:单次任务、结果校准、规则固化、对抗验证、状态修补、定时上线、运行复盘。

这次,Video Knowledge Editor 几乎完整演示了这七步为什么不能省。

我先给它一条 35 分钟的 AI 营销视频做 baseline。它在返回的报告里说,自己拿到了 720p 视频、15 个章节、YouTube 自动识别文本和本地 Whisper 转录;同时也老老实实记录了 YouTube 字幕接口的 429 限制、Whisper 的 13 个长缺口,以及时间戳不确定。

这个 baseline 已经比普通的视频总结认真得多。但它还没有达到我要的标准。

我做“科爷的数字生命”,不是为了把一条视频换一种说法再发一遍。我需要它分清三层:原视频到底说了什么,这些内容应该如何理解,以及我能从中提出什么新的视角。如果只有前两层,那不是我的文章,只是一份改写得更顺的视频摘要。

02-framework-video-three-layers-v2.png

所以第二步 calibration 里,我保留了它已经做对的部分:公开来源获取状态、转录不确定性、证据映射、真实视频截图、知识资产与文章分开、未经我审核不发布。然后再改掉几个会让系统越跑越偏的设定。

比如,它原本想为每条视频固定保留 6 到 12 张图。我把这个数量指标删了,只留能解释框架、证明画面与口播差异,或者值得未来检索的画面。它原本会把 Brand 5.0 Lab、GEO、Grok Bot 等等都当作可能的个人关联,我给这些关联加了条件:只有原始材料真的说到营销或品牌,Brand 5.0 Lab 才能出现。

个性,不是把自己的每一个项目硬塞进每一篇文章。

个性来自取舍。

专门挑一个会让它翻车的新任务

校准完成后,我让 Video Knowledge Editor 把方法保存为 Skill,但没有马上建 Routine。

我给它换了一种完全不同的材料:上一次是一个人对着投影讲框架,这一次是有多位说话者、没有官方章节的纪录片式访谈。

这次验证暴露了更隐蔽的问题。标题里的 9650 亿美元,并没有被任何人在口播里直接说出来;解说只说了“接近一万亿”,而 9650 亿出现在开头的 Bloomberg 标题画面中。另一组营收数字也只出现在画面,字幕里没有。

普通摘要很容易把标题、口播和画面全部压成一条“视频说了”。这只 Bot 没有这样做。它保留了冲突,也承认多位说话者的身份是根据上下文推断,不是字幕直接标注。

最终结果是 PASS WITH LIMITATIONS。

这个结果比一个干干净净的 PASS 更有价值。因为它证明这只 Bot 已经学会了一件很重要的事:不能核验的地方,不需要表演确定。

验证后我又做了一次状态修补:即使有人工字幕,也尽量再取一轨自动识别文本用来对照;字幕没有说话者 ID 时,必须生成带置信度的 speaker map;证据映射不能长成另一份逐字稿;纪录片画面与时间戳发生错位时,要显式记录。

这些都不是一句“写得更准确”能解决的问题。它们需要成为可执行、可检查的规则。

内容 Bot 最容易偷走的,是“作者权”

到 Idea to Article Editor 时,问题又变了。

一只只做摘要的 Bot,准确是第一标准。一只负责最终文章的 Bot,除了准确,还必须解决另一个问题:谁有资格代表“我”说话?

我给它设了明确的 source_class。X Intelligence Scout、Video Knowledge Editor、Brand 5.0 Lab Researcher 产生的材料,都是上游工作材料;只有被标记为 user-authored 的内容,才能直接建立我的观点和语气。

我可以让几只 Bot 替我找材料、处理视频、列证据,但我不想让一只上游 Bot 因为文笔太成熟,就把它的判断悄悄写成我的判断。如果连这条线都没有,“科爷的数字生命”很快就会变成“几只 Bot 的内容自动生成机器”。

03-framework-authorship-source-class.png

在对抗验证里,我故意把四种材料混在一起:我自己写的观点、视频 Bot 的产出、X 信号 Bot 的产出、Brand 5.0 Lab 的研究。它必须先判断材料是 READY、READY WITH VERIFICATION GAPS 还是 NOT READY,然后再决定能不能写。

它还要识别另一个很常见的假象:三只 Bot 都提到同一个事件,不等于我获得了三份独立证据。如果它们最后都来自同一条原始新闻,那仍然只有一个来源。在测试里,我把这叫作 thematic rhyme ≠ corroboration。

验证的结果不是满分开场,它泄露了一句编辑控制室里的话:I am not treating them as independent results。这句话说明它做了正确的证据判断,但这种自我检查的语言不应该泄露到读者看到的正文里。

我没有因为这一句话重写整个 Skill,而是把它放进 Routine 的 final audit。这是执行瑕疵,不是设计缺陷。调试 Bot 最容易犯的错,是只要看到一次失败,就把整套提示词推倒重来。其实更重要的是,先判断问题属于角色、方法,还是运行层。

Routine 上线了,也不等于完成

Idea to Article Editor 的 Routine,我没有设成“每天自动写一篇文章”。这种设定看起来很爽,但它会迫使 Bot 在材料不够时也制造一篇东西。

我给它建的是 Editorial Inbox:只处理我明确放进队列的 editorial packet,工作日晚上 9 点检查,每次最多处理一份最早的 pending 材料。如果队列是空的,它只返回“No pending packets in the Editorial Inbox.”然后停止。

遇到证据有缺口的材料,它可以生成 RESEARCH_REQUESTS.md,但不能自动去执行这些研究请求。这个边界对我很重要:上游 Bot 可以告诉我还缺什么,但是否继续研究,仍然由我决定。

创建完 Routine 后,我又做了一次完整的 acceptance test:

enqueue
  → pending
  → manual run
  → readiness gate
  → outputs
  → queue update

测试包含四种来源,最终被判定为 READY WITH VERIFICATION GAPS。Bot 生成了 ARTICLE_DRAFT.md、EDITORIAL_NOTES.md 和 RESEARCH_REQUESTS.md,没有把 Research Requests 自动派出,也没有修改上游材料。队列状态从 pending 走到 processing,再到 completed。之前泄露到正文里的控制室语言,也被 final audit 清掉了。

到这里,第七步才算真正完成。

04-framework-editorial-inbox.png

一个 Routine 存在,只能证明定时器存在。它能从真实入口读到正确数据,走完正确状态,产生预期文件,并在该停下的地方停下,才证明它可以工作。

把七步法真正用起来的 7 个提示词

下面这一套提示词,不只适用于 Grok Bot。只要你正在创建一只会使用文件、网页或自动任务的 Agent,都可以复制、粘贴使用,只需要把方括号里的内容换成自己的需求。

第一步,先建立 baseline

Complete one real baseline task for [goal].

Use these inputs only: [sources / files / websites].
Return: [expected deliverables].

Before working, state what inputs are accessible and what is missing.
Separate verified facts, source claims, inferences, and unknowns.
Preserve links or file references for every important conclusion.

Do not create a Skill or Routine yet.
Do not publish, send, purchase, delete, or modify external systems.
Stop and report any access failure or uncertainty you cannot resolve.

第二步,结果校准:把这次的修正变成长期规则

Review the baseline result against this goal: [goal].

Keep the parts that worked: [list].
Correct these problems: [list].

For every correction, write a reusable rule with:

  • the condition that triggers it

  • the required action

  • the forbidden shortcut

  • the expected evidence or output

  • what to do when the evidence is unavailable

Do not merely rewrite the previous answer.
Show the revised workflow and the remaining unknowns.
Do not create a Routine.

第三步,规则固化:保存为 Skill

Save the calibrated workflow as a Skill called [skill name].

The Skill must define:

  1. when to use it

  2. required inputs and access

  3. the exact work sequence

  4. decision and evidence rules

  5. output files and formats

  6. failure and partial-completion handling

  7. approval and safety boundaries

Treat [canonical source / path] as the source of truth.
Do not rely on chat memory when that source is unavailable.
Show the saved Skill and its location.
Do not create or run a Routine.

第四步,对抗验证:专门测它最容易犯的错

Validate [skill name] on this deliberately difficult input: [test input].

This test is designed to expose: [likely failure modes].
Use a different source type or edge case from the baseline.

Check whether the Skill:

  • preserves source disagreements

  • distinguishes evidence from inference

  • avoids duplicate or false corroboration

  • keeps user-authored views separate from upstream Bot output

  • reports access, state, and confidence limitations

  • stops at every approval boundary

Do not edit the Skill during the test.
Return PASS, PASS WITH LIMITATIONS, or FAIL, with exact evidence.

第五步,状态修补:修路径、事实源和失败处理

Using the validation report, patch only the operational weaknesses in
[skill name].

Define:

  • canonical source-of-truth locations

  • allowed status values and transitions

  • deduplication and stale-data rules

  • missing-input and access-failure behavior

  • retry and idempotency rules

  • separation between working files, reports, and final outputs

Do not change validated editorial or decision rules unless the report shows
a design defect. List each change and the validation evidence that requires it.
Do not create or run a Routine.

第六步,定时上线:最后才创建 Routine

Create one Routine owned by [Bot name] using [skill name].

Schedule: [schedule and timezone].
Input source: [queue / canonical store / website].
Maximum work per run: [limit].
Expected outputs: [deliverables and location].

If there is no new input, return [exact no-op message] and stop.
If an input is missing or stale, record the failure and preserve the item.
Require approval before any sending, publishing, purchase, deletion,
permission change, or production write.

Show the Routine ID, owner, schedule, next run, input, output, status model,
empty-input behavior, and failure behavior.
Do not run it yet.

第七步,运行复盘:用真实队列做验收

Run [Routine name] once now as a manual acceptance test.

Use the current production Routine and the saved Skill exactly.
Process only [the oldest pending item / one safe test item].
Do not modify the Skill, Routine, schedule, schema, or upstream sources.

Report:

  1. item ID and status transition

  2. inputs actually loaded

  3. Skill loaded successfully: Yes / No

  4. expected outputs created: Yes / No

  5. source and evidence checks passed: Yes / No

  6. approval boundaries preserved: Yes / No

  7. upstream or external systems modified: Yes / No

  8. queue or state updated correctly: Yes / No

  9. output location and remaining limitations

Return exactly one status: PASS, PASS WITH LIMITATIONS, or FAIL.

我要的“数字生命”,不是一个“更像我”的人

这次调试完成后,我没有继续加第四只、第五只 Bot。

原因很简单:对我而言,数字生命的价值已经不在于“它能不能像我一样说话”,而在于它能不能接住我真实生活里的工作。

我想看演出,Theatre Scout 要知道什么值得我安排时间甚至专程旅行;我要写 AI 与数字生活,X Intelligence Scout 就不能把每条热闹都丢给我;我要做 Brand 5.0 Lab,Researcher 就必须保留反证,不能顺着我的期待制造市场结论;我要写文章,上游 Bot 就可以帮我处理材料,但不能拿走最终的判断权。

这些需求看起来不够“通用”,甚至有点琐碎。但“数字生命”如果没有这些琐碎,就只剩下一张很漂亮的“系统架构图”。

七步法的本质,也不是让提示词变得更长。它是把一只 Bot 从“偶尔能做对”,一步一步推到“在指定条件下可重复地做对”。

创建一只 Bot,几分钟就够了。

把它变成数字生命的一部分,靠的还是试做、校准、验证和长期复盘。