转译:给 Claude Fable 5.1 写提示词

译自 Claude 官网 原标题:Prompting Claude Fable 5.1

作者:Claude | 原文发布:2026年9月1日

模型能力、API 变更、定价和可用性,见 Claude Fable 5.1 有什么新变化。适用于各 Claude 模型的通用技巧,见 提示词最佳实践。

你现有的 Claude Fable 5 提示词,换到 Claude Fable 5.1 上大多不用改就能用好,但有一批行为差异值得单独了解。先对上你观察到的症状:

  • 不确定该跑哪个 effort(投入档位),或延迟和成本高于任务所需:[把所有 effort 档位都试一遍](#consider-all-effort-levels)

  • 工具调用之间几乎没有文字:[要求面向用户的进度更新](#ask-for-user-facing-progress-updates)

  • 智能体循环(agent loop)里每回合只发一次工具调用:[在智能体循环里批量发出互不依赖的工具调用](#batch-independent-tool-calls-in-agent-loops)

  • 请求报错 bound to a different conversation(绑定到了另一段对话),或你的 harness(编排层)会在请求之间改动更早的回合:[对话历史只追加、不改写](#keep-the-conversation-history-append-only)

  • 行文又长又密:[行文密度](#writing-density)

  • 聊天回复的结构不够撑起内容:[聊天里的排版](#formatting-in-chat)

  • 摘要照搬了来源措辞,却没标成引文:[引用检索到的原文](#quoting-retrieved-sources)

  • 回合在工作完成前就结束了,或模型对你已经要求过的工作还在征求许可:[把整件任务做完](#finish-the-whole-task)

  • 客户端压缩(compaction)摘要丢掉了约束、决策或确切细节:[告诉模型压缩摘要必须保留什么](#tell-the-model-what-to-preserve-in-compaction-summaries)

  • 出现了没要求的修复或扩展,或提交的测试文件比任务需要的多:[改动和测试只覆盖任务要求的范围](#keep-changes-and-tests-to-what-the-task-asks-for)

  • 在 low effort 下凭记忆作答、不去搜索:[低 effort 时的搜索触发](#search-triggering-at-low-effort)

  • 无害的编程请求返回 stop_reason: "refusal"(拒绝):[减少安全护栏误报](#reduce-safeguard-false-positives)

  • 小改动却整文件重写:[优先做针对性修改,不要整文件重写](#prefer-targeted-edits-over-whole-file-rewrites)

  • xhigh 或 max effort 下的长交付物耗时过久,或撞上 max_tokens:[在 xhigh 和 max 档为长输出留足空间](#leave-room-for-long-outputs-at-xhigh-and-max-effort)

  • 子智能体在跑时主智能体闲着:[子智能体运行时,让主智能体继续干活](#let-the-lead-agent-keep-working-while-subagents-run)

  • 关于图表和信息密集图像的回答漏掉细节:[给视觉任务提供裁剪和放大工具](#give-vision-work-tools-to-crop-and-zoom)

把所有 effort 档位都试一遍

先从默认 effort 档位 high 起步,再用你自己的评测去试其余档位(low、medium、xhigh 和 max)。在 Claude Fable 5.1 上,effort 是权衡智力、延迟和成本的主旋钮。即便你已经在 Claude Fable 5 上扫过一遍,也请再扫一次:档位名称在不同模型之间并不对应同样多的思考量。

Claude Fable 5.1 相对 Claude Fable 5 的能力提升,在各档 effort 上都能看到,且在更高档位最明显。在 medium 上,结果大致对得上 Claude Fable 5,成本更低;所以评测显示质量还站得住的地方,就降到 medium 或 low。在 low 上,Claude Fable 5.1 的单任务成本往往能跟 Claude Opus、Claude Sonnet 竞争,分数还更高;因此凡是你本来会用更小模型、更高 effort 的场景,都把它放进对比。

有两类跟档位绑定的行为单独成节:在 low 上,Claude Fable 5.1 更少调用搜索和检索工具(见 [低 effort 时的搜索触发](#search-triggering-at-low-effort));在 xhigh 和 max 上,它在写出长交付物之前可以思考更久(见 [在 xhigh 和 max 档为长输出留足空间](#leave-room-for-long-outputs-at-xhigh-and-max-effort))。

要求面向用户的进度更新

Claude Fable 5.1 的默认行为是:在长工具调用回合里,面向用户的进度更新比 Claude Fable 5 更少。effort 越高、工具链越长,这一点越明显。用户会看到智能体连续几分钟没有动静,或最后一条消息只覆盖了最后一步,而不是整项任务。

先确认你的客户端到底有没有收到进度更新。模型在工具调用之间写下的短注——刚发现了什么、接下来要做什么——会以 进度更新 thinking 块 的形式返回;而默认的 thinking.display 是 "omitted",这些块是空的。把 display 设为 "updates"(beta,thinking-display-updates-2026-08-18 请求头),并把每个非空的 thinking 块渲染成状态行;或者设成 "summarized",就会连同摘要后的推理一起返回这些更新。如果你根本没去请求这些块,模型的更新可能只是没送到用户眼前。

第二,检查你的提示词里有没有压住叙述的指令。更早的一些模型干活时很爱播报,于是系统提示里出现了「hold all findings for the final response」这类句子。先删掉这类句子,再考虑往上加。

如果你还想要更多更新——比如结对编程,或其他人在回路(human-in-the-loop)的工作——加一句短的系统提示,写明你希望模型何时给出面向用户的文字,以及每条更新该包含什么:

Before you start, say in a line what you're about to do; brief updates while you work help the user follow along. Close with a short recap that stands on its own — what you found, what you did, and what's next — so a reader who only sees the last message has the full picture.

如果你的产品会折叠或隐藏工具输出,要告诉模型。否则它可能跑一些命令来「展示」输出给用户,而你的界面根本不显示那些内容。把这条说明放进一条 回合作用域系统消息(clearat: "nextuser_message",beta):

Only you see that command's output — the user's terminal shows at most a few lines of it. If the user needs to read any of it, put it in your reply.

在智能体循环里批量发出互不依赖的工具调用

Claude Fable 5.1 通常会按预期发出并行工具调用:请求点名要取好几样东西时,它会并行发出这些调用。例外出在编程和 computer use(计算机操作)循环——下一步互不依赖的调用是任务暗示的,而不是明确点名的(自定义编程智能体、bash 加编辑器的 harness、computer use):在这类循环里,它可能改成每回合只发一次。这不影响回答质量,但多出来的每个回合都要消耗 token、一次往返和实际耗时。在当前请求末尾加一句提醒就能解决:

First privately list what you need next; then request every item that doesn't depend on another's result in this one response.

每次把工具结果发回去时,把它追加到那条用户消息后面,写成一条 回合作用域系统消息:messages 里一条 role: "system" 的记录,并带上 clearat: "nextusermessage"。一旦后面出现了新的用户消息,API 就会清掉更早的副本,模型因此只读到最新那一条。回合作用域系统消息仍是 beta,需要带上 beta 请求头 mid-conversation-system-clear-at-2026-08-21。没有 beta 的话,就把这句话放进同一条用户消息里、紧接在 toolresult 块后面的文本块中。

每一回合都追加一份新副本,更早的副本留在原地,保持字节级完全一致。它们还在数组里,但一旦被清除,模型就看不见,也不消耗输入 token。删掉或改写它们,等于在改更早的回合:会从那一点重启 提示缓存,并作废其后出现的 thinking 块(见 [对话历史只追加、不改写](#keep-the-conversation-history-append-only))。

对话历史只追加、不改写

历史要只追加(append-only):每一轮助手回复都按 API 返回的原样接上去,thinking 块也算在内;请求与请求之间,不要改写更早的回合。对 2026 年 8 月 31 日及之后新建的账号,Claude Fable 5.1 的 thinking 块只在产生它们的那次对话里有效:前缀(系统提示、工具列表,或任何更早的消息)一旦改过,再回放 thinking 块,请求会返回 400;若设置 thinking.blockbinding.prefixmismatchbehavior: "dropblock"(beta,请求头 thinking-binding-controls-2026-08-01),则会丢弃受影响的块。后续模型预计会对所有账号强制这项检查,所以即便你现在的账号还没被强制,也该现在就按这个模式来接。

会触发这项检查的历史改写,也正是会重启提示缓存的那些操作:每轮把提醒插进去再删掉、把更早的回合就地改成摘要,或在会话中途改系统提示。每轮提醒用回合作用域系统消息发送;要改指令或工具,发一条对话中途系统消息,不要去改写 system 或 tools。需要裁剪时,交给服务端的压缩(compaction)或上下文编辑。如果在客户端做压缩,最简单的形态是:整段历史换成一条摘要消息加上新的用户回合,其余一律不回放——thinking 块不带过去,就不会失败,模型会在压缩后的对话上重新思考(见在客户端自定义压缩)。缓存读取现在更便宜了(见定价),为省钱而过早压缩,在 Claude Fable 5.1 上未必仍是成本和智力之间的正确取舍,所以不妨把压缩点往后挪,自己试一试。

要找出 harness 已经在做的改写,用 prefixmismatchbehavior: "dropblock" 跑一轮会话,并记录 inputtransformations,做法见如何判断你的集成是否受影响;或者抓取它在若干正常回合里发出的原始请求,确认相邻请求除了新追加的回合之外,其余部分字节级完全一致。

行文密度

Claude Fable 5.1 的文笔总体上比前几代 Claude 更好:套话更少,未加解释的行话也更少。不过有时行文会比 Claude Fable 5 更密:句子更长,分段更少。加一条指令,把反面模式——矫揉造作的文风(mannered prose)——定义清楚,会有帮助。它针对的是 "a dial worth turning"、"this point earns its keep" 这类卖弄比喻、并不直说的句子。优先写进用户消息,也可以放进系统提示:

Mannered prose substitutes metaphor and flourish for direct statement. Instead of "a parameter worth varying," the mannered writer produces "a dial worth turning." Instead of "this point still matters," they write "this point earns its keep." The phrases exist to display the writer, not to convey the idea, and readers can tell. That is why mannered prose irritates: it makes the reader work harder so the writer can perform. It is also imprecise. Metaphors drag in connotations the writer did not choose and cannot control. The fix is to say what you mean. When a literal phrase is available, use it.

短版往往也管用:

Please remove all mannered prose.

聊天里的排版

更早的模型在聊天里会把项目符号和加粗用过头,不少提示词因此带着反排版规则,专门用来压住这种习惯。Claude Fable 5.1 偏向另一头:加粗用得更少,也不那么爱用标题、列表或引号。如果你的提示词里有反排版的措辞,删掉,或换成一条说明何时该用特定排版的规则,例如:

Use lists and bullet points when asked to, or when the content is multifaceted enough that they help with clarity. If the person explicitly requests minimal formatting, always format your responses without bullet points, headers, lists, or bold emphasis, as requested. In conversational, personal, or emotional exchanges, keep to plain prose.

引用检索到的原文

做文档摘要时,Claude Fable 5.1 比 Claude Fable 5 更容易把原文段落复述出来,却不标成引文。应对办法是:在系统提示里加一个完整的正确回复示例——用户请求、回复,再加一句说明为何正确。

<example>
<user>look up how the Riverton Ledger and the Coast Dispatch each covered the Harbor Bridge closure and compare their reporting</user>
<response>
[web_search: Harbor Bridge closure Riverton Ledger]
[web_search: Harbor Bridge closure Coast Dispatch]
Both outlets agree on the basics: the bridge closed on March 3 after inspectors found cracked welds, and the state expects repairs to take about eight months. Where they differ is emphasis. The Ledger treats it as a local-economy story. The Dispatch frames it as a funding failure; its editorial calls the closure "entirely foreseeable." Read together, the Ledger explains who is affected now and the Dispatch explains how it came to this — neither account alone gives the whole picture.
</response>
<rationale>CORRECT: The response is organized around where the two outlets agree and differ, not as a walk through either article. Each outlet's reporting is conveyed in one or two sentences of the assistant's own indirect speech. One short marked phrase from one source; every other claim is reworded. The response is still specific and complete.</rationale>
</example>

把那两行 [web_search: ...] 换成你自己的工具名,让模型把它们读成模板化的工具输出,而不是要原样打出来的文字。

把整件任务做完

目标清楚时,Claude Fable 5.1 几乎不需要方法上的指导,就能执行很长的任务。但在复杂的异步工作负载上,要推它一把:不要在工作做完之前就结束回合。没有这句提醒,模型有时会只描述下一步要做什么而不去做("Next, I'll …"),或停下来,为原始请求已经覆盖的步骤征求许可("Shall I apply this?")。用户就得回 "continue" 或 "go ahead"。这对结对编程和其他人在回路的工作合适,但没有用上模型完整的长程能力。

两段系统提示加在一起可以缓解这个问题。两段都要用。如果必须控制提示长度,只用第一段,大部分效果还在。第一段告诉模型:已经要求过的工作不要再问,自己说出的下一步要去做:

You are operating autonomously. The user is not watching in real time and cannot answer questions mid-task, so asking 'Want me to…?' or 'Shall I…?' will block the work. For reversible actions that follow from the original request, proceed without asking. Stop only for destructive actions or genuine scope changes the user must decide. Offering follow-ups after the task is done is fine; asking permission before doing the work is not.

Exception: when the user is describing a problem, asking a question, or thinking out loud rather than requesting a change, the deliverable is your assessment. Report your findings and stop. Don't apply a fix until they ask for one.

Before ending your turn, check your last paragraph. If it is a plan, an analysis, a question, a list of next steps, or a promise about work you have not done ('I'll…', 'let me know when…'), do that work now with tool calls. That includes retrying after errors and gathering missing information yourself. Do not stop because the context or session is long. End your turn only when the task is complete or you are blocked on input only the user can provide.

Before running a command that changes system state (such as restarts, deletes, or config edits), check that the evidence actually supports that specific action. A signal that pattern-matches to a known failure may have a different cause.

开篇那句——告诉模型用户并不在实时看着——贡献了大部分效果。按原文保留。如果你的产品需要模型在特定确认点停下,在这句后面加一句,把要确认的事项列出来。这段也会让模型更少就含糊的请求发问,所以在你自己的任务上核对一下这个取舍。

第二段把用户请求定义为交付物的范围:

# Delivering work
The user's request — or the plan they approved — sets the scope, and the scope is the deliverable: don't quietly narrow, widen, or swap it. Read ambiguity the way a careful colleague would: make routine judgment calls yourself, and check in only when different readings would lead to materially different work. If you see a real problem with the task as specified, say so in a sentence or two and keep building under stated assumptions; if the user hears the concern and reaffirms, that is their decision, so deliver the full request.

If a question comes up partway, first do everything that doesn't depend on the answer; then state the assumption you made, or — when going ahead on a wrong guess would be unsafe or would make the work useless — put the question at the end of a turn that also delivers that progress. If one part turns out to be blocked, complete every other part in full and say exactly what you left out and why — the whole task is the deliverable, and scaling it down is the user's call, not yours. A step you have decided on is something to run, not to announce: describing the next step and ending the turn leaves it undone until the user replies.

Keep changes to what the request needs. Something else you notice worth doing — cleanup or documentation the task didn't call for, a change to a file the task didn't require — is a suggestion to make at the end, not a change to make; actions clearly beyond what the ask implies, and risky or destructive ones, still need the user's go-ahead.

告诉模型压缩摘要必须保留什么

长对话被压缩时,明确告诉 Claude Fable 5.1 摘要必须保留什么,它会按你的要求来。服务端 压缩 已经会这样做。如果你在客户端做压缩,用下面这条摘要指令:

Summarize the transcript inside <summary></summary> tags. Include relevant information in the summary such that this conversation will be continued by a new context window without needing to redo work or be reprovided with relevant constraints or context. Be sure to preserve: (1) any difficulties or problems that came up, and how they were handled or resolved; (2) any possibilities, options, or approaches that were raised, tried, or set aside, and why; (3) anything that was asked for, decided, agreed, ruled out, or established as a preference, constraint, or boundary — stated exactly; (4) exactly where things stand now — what has been covered, settled, or completed so far; (5) anything still open, unresolved, promised, or expected to happen next; (6) specific details that would be hard to reconstruct — names, numbers, dates, exact wording, links or references — kept exactly. Be complete on these even at the cost of length; keep everything else concise. Weight the two voices differently: keep what the user said, asked for, shared, or established carefully and close to their own words; your own explanations and reasoning can be condensed much further, to what they concluded or produced — as long as nothing in the six items above is dropped.

改动和测试只覆盖任务要求的范围

让它实现开放式功能时,Claude Fable 5.1 会把要求的做完,有时还会多做:修旁边的代码、扩展任务没提的行为,或提交超出这次改动该有的测试文件。明确写出哪些不要动,它会照做。加上下面这条指令后,未经要求的加量和提交进仓库的测试代码会明显下降,任务成功率没有可测到的变化:

If, while working or testing, you find a pre-existing bug, a performance concern, or behavior the task doesn't mention, don't fix, optimize or extend it in this change unless the requested behavior cannot work without it; report it as a follow-up in your summary. Where the task is ambiguous, implement the reading its wording and the surrounding code most directly support, state that assumption in your summary, and don't build for the other readings as well. Verify your work however you like; scratch scripts and quick checks need not be kept. Commit tests only where the task asks for them or this repository already keeps tests for this kind of change, sized like the neighboring test files — roughly one focused test per stated behavior — and don't turn scratch checks into additional permanent test files. This is about extras only: implement every behavior the task asks for, completely.

低 effort 时的搜索触发

在 low effort 下,Claude Fable 5.1 比 Claude Fable 5 更少调用搜索或检索工具,更常凭记忆作答。有时最简单的修法是只把受影响的那几轮调高 effort,而不是整段对话一起抬。见 在对话中途更改 effort。

另一些情况,提示词里朝「先核实」轻轻推一把就够。在系统提示里写清楚:认出一个名字,不等于知道它现在的状态;这类名字应按用户写下的原文去搜:

When a query centers on a name you do not confidently recognize, or recognize from a fast-moving area like AI models and developer tools where the landscape shifts within months, the name itself is the thing to verify: search before answering, and include the name as the user wrote it in at least one query alongside any reformulations. This holds even when you have some background on it — partial background is exactly what makes an out-of-date answer sound authoritative, so familiarity is not a reason to skip the search.

减少安全护栏误报

Claude Fable 5.1 的安全分类器产生的误报,比 Claude Fable 5 刚上线时更少,而且允许在源代码里找漏洞。误报仍会发生,被拦下的请求会返回 stop_reason: "refusal"(见 拒绝、回退与计费)。下面三种情况更容易触发:

  • 编译检查的问法: 不要问 "Does this program compile without errors?",改问 "Are there any bugs in this program?"

  • 小众编程语言: 给模型补上这门语言是什么、怎么工作,比如让它能读到该语言的文档。

  • 工具输出里的 Base64: 把 base64 编码数据送进模型上下文的工具可能触发误报,建议的修法是去掉这类输出。

优先做针对性修改,不要整文件重写

如果 Claude Fable 5.1 为小改动整文件重写,把下面这条指令追加到系统提示或第一条用户消息。相比 Claude Fable 5,Claude Fable 5.1 更常重写整个文本文件,而不是做针对性修改。改完的文件通常一样,但除非文件很短、或大部分内容都在改,重写会多花输出 token 和时间。这条指令能让 Claude Fable 5.1 在中小改动上回到和 Claude Fable 5 一致。

The number of tokens used to edit files is best minimized, all else being equal. Therefore, when it will not affect the end result, try to surgically edit a file rather than rewrite the entire thing.

在 xhigh 和 max 档为长输出留足空间

在 xhigh、尤其是 max effort 下,Claude Fable 5.1 可能思考更久才开始写回复。如果一次请求要一份很长的交付物,比如整篇长文档重写,它可能先在 thinking 里把交付物起草大半,再在回复里写一遍,等待更久,输出 token 也更多。最简单的做法是这类请求先跑 high(推荐起点),只有你测到质量提升时,再升到 xhigh 或 max(见 把所有 effort 档位都试一遍)。如果确实要在 xhigh 或 max 上跑:

  • 把 max_tokens 设得够 thinking 和回复一起用,而不是只按你预期的回复长度来设。

  • 把下面这段说明追加到用户消息末尾。对文案和代码请求,它会让 thinking 短很多。把 [maxtokens] 换成这次请求实际的 maxtokens 值,例如 64,000。

Everything produced in one reply, including any reasoning or drafting it does before the reply, counts toward a single limit of about [max_tokens] tokens. If that limit is reached before the reply is finished, the person receives a cut-off response and has to start over. Composing an entire output or deliverable in full as reasoning and then again as a reply would double the length of the turn without improving the result, so don't do that.

Instead, when the person has asked for a long or effort-intensive deliverable such as a multi-section document, a large table or dataset, or a complete code file, spend extra effort on understanding the request, checking the inputs the answer depends on, settling the structure and other difficult decisions, and otherwise using the reasoning space to reason and the output space to write an output. Usually it is not needed to draft an output multiple times.

子智能体运行时,让主智能体继续干活

如果你的编码智能体允许 Claude Fable 5.1 把工作委派给子智能体(subagent),不要强迫主智能体(lead agent)停下来等每一个。在编码任务上,让主智能体在子智能体运行时继续干活,能在质量、token 用量和成本相近的情况下,降低平均完成时间。做法是:

  • 启动子智能体的工具应立即返回。

  • 每个子智能体的结果就绪后,用后续一条 user 消息交回主智能体。

  • 给主智能体另备一个工具,它想等结果时再调用。

模型仍然经常选择等待。省下来的时间,来自它接着去干别的活的那些回合。

给视觉任务提供裁剪和放大工具

Claude Fable 5.1 开箱即用的视觉能力更强;面对密集图表这类复杂视觉输入,它能迭代分析、裁剪并目视核对所见时,发挥最好。要把这部分能力用满,把模型当智能体来跑,并给它一个容器:里面放原始图片或视频,并预装基础图像处理库(如 PIL 和 OpenCV)。如果跑容器开销太大,单一个图片裁剪工具就能拿到大部分提升:这个工具返回图像中选定区域的裁剪放大版,让模型把特定细节看得更细,测试时计算(test-time compute)也会随图像 token 增加而上去。一份可落地的定义见 裁剪工具示例。