# PostHog 將 Agent context 當程式碼資產持續清理並以 regression eval 驗證

> 📖 本站完整內容索引（documentation index）：[llms.txt](/llms.txt)

> 原作者：PostHog (@posthog) · 策展與摘要：EasyVibeCoding · 平台：X (Twitter) · 熱度：🔥🔥🔥 · 日期：2026-09-01

> 原始來源：https://x.com/posthog/status/2094485724171223409

## 證據與延伸閱讀

- [PostHog 將 Agent context 當程式碼資產持續清理並以 regression eval 驗證。](https://posthog.com/newsletter/context-engineering) — 一手來源 · 支持主張：The article adds manual review after /doctor, failure prompts as regression evals, context-mill sample apps and PR grading, and a loop that clusters and verifies agent feedback before subagent reproduction or fixes.
- [Model guidance | OpenAI API](https://developers.openai.com/api/docs/guides/latest-model?model=gpt-5.6) — 一手來源 · 最後核對：2026-09-01 · 支持主張：The article adds manual review after /doctor, failure prompts as regression evals, context-mill sample apps and PR grading, and a loop that clusters and verifies agent feedback before subagent reproduction or fixes.
- [PostHog context-as-code practices — github.com](https://github.com/PostHog/wizard/pull/884) — 官方 Repository · 最後核對：2026-09-01 · 支持主張：The article adds manual review after /doctor, failure prompts as regression evals, context-mill sample apps and PR grading, and a loop that clusters and verifies agent feedback before subagent reproduction or fixes.
- [PostHog context-as-code practices — code.claude.com](https://code.claude.com/docs/en/commands) — 一手來源 · 最後核對：2026-09-01 · 支持主張：The article adds manual review after /doctor, failure prompts as regression evals, context-mill sample apps and PR grading, and a loop that clusters and verifies agent feedback before subagent reproduction or fixes.
- [PostHog context-as-code practices — github.com](https://github.com/PostHog/posthog/pull/75873) — 官方 Repository · 最後核對：2026-09-01 · 支持主張：The article adds manual review after /doctor, failure prompts as regression evals, context-mill sample apps and PR grading, and a loop that clusters and verifies agent feedback before subagent reproduction or fixes.
- [PostHog context-as-code practices — github.com](https://github.com/PostHog/context-mill/blob/main/context/commandments.yaml) — 官方 Repository · 最後核對：2026-09-01 · 支持主張：The article adds manual review after /doctor, failure prompts as regression evals, context-mill sample apps and PR grading, and a loop that clusters and verifies agent feedback before subagent reproduction or fixes.
- [PostHog context-as-code practices — github.com](https://github.com/PostHog/wizard-workbench/tree/main/services/wizard-ci) — 官方 Repository · 最後核對：2026-09-01 · 支持主張：The article adds manual review after /doctor, failure prompts as regression evals, context-mill sample apps and PR grading, and a loop that clusters and verifies agent feedback before subagent reproduction or fixes.
- [PostHog context-as-code practices — github.com](https://github.com/PostHog/wizard-workbench/tree/main/services/pr-evaluator) — 官方 Repository · 最後核對：2026-09-01 · 支持主張：The article adds manual review after /doctor, failure prompts as regression evals, context-mill sample apps and PR grading, and a loop that clusters and verifies agent feedback before subagent reproduction or fixes.
- [PostHog context-as-code practices — github.com](https://github.com/PostHog/wizard-workbench/pulls?q=is%3Apr+is%3Aclosed+label%3ACI%2FCD) — 官方 Repository · 最後核對：2026-09-01 · 支持主張：The article adds manual review after /doctor, failure prompts as regression evals, context-mill sample apps and PR grading, and a loop that clusters and verifies agent feedback before subagent reproduction or fixes.
- [PostHog context-as-code practices — github.com](https://github.com/PostHog/wizard-workbench) — 官方 Repository · 最後核對：2026-09-01 · 支持主張：The article adds manual review after /doctor, failure prompts as regression evals, context-mill sample apps and PR grading, and a loop that clusters and verifies agent feedback before subagent reproduction or fixes.
- [PostHog context-as-code practices — github.com](https://github.com/PostHog/wizard/blob/ba91f26c418f332f1ede8b6e80fd9fa14bbb22e8/src/lib/agent/runner/sequence/orchestrator/queue-tools.ts) — 官方 Repository · 最後核對：2026-09-01 · 支持主張：The article adds manual review after /doctor, failure prompts as regression evals, context-mill sample apps and PR grading, and a loop that clusters and verifies agent feedback before subagent reproduction or fixes.
- [PostHog context-as-code practices — github.com](https://github.com/PostHog/context-mill/pull/272) — 官方 Repository · 最後核對：2026-09-01 · 支持主張：The article adds manual review after /doctor, failure prompts as regression evals, context-mill sample apps and PR grading, and a loop that clusters and verifies agent feedback before subagent reproduction or fixes.
- [PostHog context-as-code practices — github.com](https://github.com/PostHog/posthog/blob/27c2aeee512923e4c9b29045d84a7ec932312ebf/.github/pull_request_template.md?plain=1) — 官方 Repository · 最後核對：2026-09-01 · 支持主張：The article adds manual review after /doctor, failure prompts as regression evals, context-mill sample apps and PR grading, and a loop that clusters and verifies agent feedback before subagent reproduction or fixes.
- [PostHog context-as-code practices — github.com](https://github.com/PostHog/posthog/pull/81018) — 官方 Repository · 最後核對：2026-09-01 · 支持主張：The article adds manual review after /doctor, failure prompts as regression evals, context-mill sample apps and PR grading, and a loop that clusters and verifies agent feedback before subagent reproduction or fixes.
- [PostHog context-as-code practices — claude.com（貼文明示來源）](https://claude.com/blog/the-new-rules-of-context-engineering-for-claude-5-generation-models) — 一手來源 · 支持主張：The article adds manual review after /doctor, failure prompts as regression evals, context-mill sample apps and PR grading, and a loop that clusters and verifies agent feedback before subagent reproduction or fixes.
- [PostHog context-as-code practices — youtube.com（貼文明示來源）](https://youtube.com/watch?v=qyPCVqFUyDo) — 一手來源 · 支持主張：The article adds manual review after /doctor, failure prompts as regression evals, context-mill sample apps and PR grading, and a loop that clusters and verifies agent feedback before subagent reproduction or fixes.
- [PostHog context-as-code practices — posthog.com（貼文明示來源）](https://posthog.com/newsletter/agent-autonomy) — 一手來源 · 支持主張：The article adds manual review after /doctor, failure prompts as regression evals, context-mill sample apps and PR grading, and a loop that clusters and verifies agent feedback before subagent reproduction or fixes.
- [PostHog context-as-code practices — posthog.com（貼文明示來源）](https://posthog.com/) — 一手來源 · 支持主張：The article adds manual review after /doctor, failure prompts as regression evals, context-mill sample apps and PR grading, and a loop that clusters and verifies agent feedback before subagent reproduction or fixes.
- [PostHog context-as-code practices — posthog.com（貼文明示來源）](https://posthog.com/newsletter/loops) — 一手來源 · 支持主張：The article adds manual review after /doctor, failure prompts as regression evals, context-mill sample apps and PR grading, and a loop that clusters and verifies agent feedback before subagent reproduction or fixes.
- [PostHog context-as-code practices — posthog.com（貼文明示來源）](https://posthog.com/self-driving) — 一手來源 · 支持主張：The article adds manual review after /doctor, failure prompts as regression evals, context-mill sample apps and PR grading, and a loop that clusters and verifies agent feedback before subagent reproduction or fixes.
- [PostHog context-as-code practices — newsletter.posthog.com（貼文明示來源）](https://newsletter.posthog.com/) — 一手來源 · 支持主張：The article adds manual review after /doctor, failure prompts as regression evals, context-mill sample apps and PR grading, and a loop that clusters and verifies agent feedback before subagent reproduction or fixes.
- [PostHog 關於 context engineering 的重點改變](https://x.com/posthog/status/2094485724171223409)
- [PostHog Wizard 生產環境反饋分類](https://pbs.twimg.com/media/HREGBPJa8AE0or9.png)

## 中文摘要

PostHog 將 Agent context 當程式碼資產持續清理並以 regression eval 驗證。

![](https://pub-75d4fe1e4e80421b9ecb1245a7ae0d1a.r2.dev/curated/9e0b6d6900c41856.png)
> 網格背景的簡報圖解，左側標題說明 Context engineering 用於補充 base models 所缺乏的資訊，右側以一個大圓與內部不規則圖形對比「Context you provided」與「What the model knows」，右上角則有一個橘色的 build mode 按鈕。

**核心觀點** PostHog 於 2026-08-31 發表的 implementation report 指出，context engineering 的重點正在改變：過去是補上 base model 缺少的資訊，現在更重要的是刪除會妨礙 Agent 判斷的內容。隨著模型能力提升，規則、重複指示與過時文件可能反而降低表現。Anthropic 已移除 Claude Code system prompt 的 80%，Boris Cherny（Claude Code creator）建議每 6 個月刪除一次 `CLAUDE.md`；Theo 

![](https://pub-75d4fe1e4e80421b9ecb1245a7ae0d1a.r2.dev/curated/a47fdb18d46f6b4a.jpg)
> 發文者 Theo - t3.gg (@theo) 的社群貼文截圖，內容提及花費數小時手動編寫與審查 CLAUDE.md、AGENTS.md 及多個 skills 的經驗與心得。

 也曾回報，手動重寫 `AGENTS.md` 值得。作者 Jina Yoon 將新的 best practices 歸納為 judgment、interfaces 與 progressive disclosure，而不是持續堆疊 rules 與 repetition。相關原始資料見 PostHog 的 [X 貼文](https://x.com/posthog/status/2094485724171223409)，貼文於 `Sep 1, 2026` `2:01 AM` 發布，畫面顯示 `6`、`17`、`317`、`26K`、`26.3K Views` 與 `674`。 

![](https://pub-75d4fe1e4e80421b9ecb1245a7ae0d1a.r2.dev/curated/4295e7a17a04d31b.png)
> 說明 context engineering 原理的圖表，在藍色格線背景左側標示「As models improve, context engineering becomes more about subtracting as much as possible.」，右上角有橘色「build mode」按鈕，中央為交疊的圖形分別標示「Context you provided」、「What the old model knew」與「What the model knows now」，並透過指線分別指示「Keep this」與「Delete this」。

**第一層：先清理，再人工審查** PostHog 的做法不是把 `/doctor` 報告直接當成結論，而是先讓工具檢查 context，再由人逐行判斷。

- Anthropic 版本會執行基本 health checks、刪除 redundant prompts、找出 broken settings 與 unused plugins，並為 lazy loading 最佳化；它也會報告檔案使用頻率與可減量。
- PostHog 範例中，`posthog.com/AGENTS.md`（`CLAUDE.md` symlink）屬於 project scope、checked in、always loaded，估計 resident tokens 約 `~1,780`，可 trim `~350`。
- `/doctor` 建議關閉 `3 unused plugins` 與 `3 skills`，估計每個 session 平均節省 `6K tokens`；但這只是 PostHog 內部觀察，不是普遍保證。工具不檢查 instructions 的 correctness，也無法從 code 發現只存在於 GitHub settings 的狀態。

PostHog 的 merge queue 案例說明人工審查不可省略。團隊加入以下規則後，數日後暫停 queue 修 failing tests，卻忘了更新指示，導致 Agent 接受錯誤資訊 `21 hours`；一名 engineer 的 PR 卡住 `10 hours`，另一人花 `45 minutes` 調查。`claude doctor` 也抓不到，因為 queue state 位於 GitHub setting。升級後應執行並逐行人工檢查；若說不出某行能防止哪一種 failure，就刪除它。

```text
claude doctor
```

```text
All merges into master go through the Trunk merge queue. Never run gh pr merge.
```

**第二層：把失敗變成 regression eval** 每次為修正 Agent mistake 而改動 context，都應保存原始 failure prompt，形成類似測試 code 的 regression eval。實務上可將失敗 prompt 貼入 `failures.md`，每次編輯或刪除 `AGENTS.md` 內容後重跑，作為最高成本 context 的 quick test suite。

PostHog 以約 `~40 clean sample apps` 執行 wizard-ci：每個 app 建立一個 PR，但 PR 不會 merge；第二個 Agent `pr-evaluator` 依 diffs 與 session logs 評分，再把 metrics 與 reports 留在 trigger PR 和 wizard-ci PR。 

![](https://pub-75d4fe1e4e80421b9ecb1245a7ae0d1a.r2.dev/curated/7a7dafcb49364610.png)
> wizard-ci 測試管線架構圖以方塊圖與格線背景呈現 PostHog 的 context-as-code 評估流程，上方包含 context-mill、wizard 與 posthog 三大模組，向下匯流至 wizard-ci 進行自動化執行與 PR 建立，並分流至包含 android、angular、django、flask、next.js 的 test apps 以及進行評分的 pr-evaluator，透過虛線將回饋循環導回 context 與 harness。

 這能抓出模型誤判 integration 已完成、因而完全略過安裝 PostHog 等問題。不過完整 wizard-ci harness 與 evaluator thresholds 未提供，token savings 與時間縮短也不能視為一般保證。

framework-specific commandments 會與 context-mill 指示共同測試，例如：

- `15.3+` 版本在 `instrumentation-client.ts` 初始化 PostHog。
- Phoenix 或 Plug app 要在 router 前加入 `PostHog.Integrations.Plug`。
- `posthog-rs` 是 Rust SDK crate，使用 `cargo add posthog-rs`，再以 `posthog_rs::client(options).await` 建立 client。

```bash
cargo add posthog-rs
```

`context-mill/context/commandments.yaml` 是 tag-based rules 設定，共 `337 lines`、`299 loc`、`38.6 KB`；各 skill 依 tags 收集 commandments，例如 `javascript_web` 與 `javascript_node`。檔案位於 context-mill/context/commandments.yaml。

**第三層：用 Agent 回饋驅動 context 更新** PostHog Wizard 的 final instruction 要求 production 發生錯誤時回報：哪些資訊或 guidance 若放進 integration prompt 或 documentation，能避免 tool failures、erroneous edits 或其他 wasted turns。這些回饋會送回 context-mill，形成 self-driving loop，但不會直接採信 Agent： 

![](https://pub-75d4fe1e4e80421b9ecb1245a7ae0d1a.r2.dev/curated/d1b3685c053d1f5b.jpg)
> PostHog Wizard 生產環境運行中由 Agent 提供之反饋分類顯示，other 與 clean run 為次數最多的前兩大議題類別，其後依次為 notebook create schema/tool、complete_task omitted、python Posthog() vs api_key 與 products-enable missing。

1. 先將回饋 cluster，並 verify underlying issues。
2. 再由 subagents 驗證 reproduction。
3. 最後才嘗試 fix。

某月的 `other` 類別中，一半是「succeeded on the first attempt」或「posthog-js already installed」等 confirmation，另一半是未達 clustering 門檻的 long tail。形成 meaningful cluster 後，subagent 曾依 legacy instructions 重現 notebook create schema/tool 的 hand-splicing text failure，再處理對應 PR。小型 workflow 則可要求 Agent 記錄實際使用的 context files；PostHog monorepo PR template 會要求列出 invoked skills，其他 Agent 曾藉此發現並修正 skill inconsistencies，包括解除 stalled ClickHouse cleanup PRs 時找到的問題。

**規則要能對應真實系統** PostHog 的 commandments 不只描述 API，也限制容易造成 silent failure 的實作方式：

- PostHog configuration 的 keys 必須 optional；初始化與 capture 要先 guard，沒有 PostHog environment 時 build 與 boot 仍可運作。development／debug builds 不得 silent，production 則維持 no-op。錯誤格式為：

```text
<VAR> variable required by PostHog is missing or un-configured, this causes events to be silently missed. This error stops appearing once <VAR> is configured
```

- React feature flags 使用 `useFeatureFlagEnabled()` 或 `useFeatureFlagPayload()`；capture 放在 user action 的 event handler，而不是以 `useEffect` 監看 state。Next.js `15.3+` 可在 `instrumentation-client.ts` 初始化；Server Components 與 Route Handlers 則使用 `posthog-node`、`getAllFlags()` 或 `getFeatureFlag()`，並 `await posthog.shutdown()`。
- 短生命週期 Route Handlers、Server Actions 與 SvelteKit server-side capture 必須使用 `flushAt: 1`、`flushInterval: 0`，capture 後、return 前 `await posthog.flush()`，避免 freeze 或 silent drop。SvelteKit 的 `svelte.config.js` 還要將 `paths.relative` 設為 `false`。
- `posthog.capture()` 的 event properties 絕不能含 PII，例如 email、full name、phone number、physical address、IP address 或 user-generated content；這些資料應放在 `posthog.identify()` 的 person properties。自動化瀏覽器測試中，bot filter 可能 silent drop every capture，測試前要處理 `navigator.webdriver`、user agent 與 `navigator.userAgentData`，並用 `?__posthog_debug=true` 檢查 `"likely bot"`。
- PostHog SQL 查詢 ALWAYS 要有 time range filter；例如：

```sql
timestamp >= now() - INTERVAL 7 DAY
```

unique values 優先使用 `uniq()`，大型資料使用 timestamp-based pagination 而非 `OFFSET`，並先在 PostHog SQL editor 測試。

**已發布的 monorepo detection** PostHog/wizard PR [#884](https://github.com/PostHog/wizard/pull/884) 由 gewenyu99 於 **2026 年 7 月 16 日**合併，包含 **22 commits**，最終 commit `b153327`、**16 checks passed**，並於同日發布 **2.45.0（#910）**，Reviewer `sarahxsanders` approved。它讓 headless 與 CI `basic-integration` runs 先 agent-scan repository，再依 scan 選擇 app，避免 monorepo 固定使用 root 而選錯 project。

scan report 會列出每個 project 的 path、framework、是否為 supported target、是否已有 PostHog，以及恰好一個 flagged `recommended`。有 recommended supported project 時優先選它，即使該 project 已有 PostHog；否則選第一個可 instrument 且沒有 PostHog 的 project。此功能由 `basic-integration-agentic-detection` 控制，**default off**；error、超時 **60s** 或找不到 project 時，會保留 `session.installDir`，精確 fallback 到舊的 root detection。它只作用於 non-interactive `ciPreRun`，interactive runs 不變；interactive picker 仍不是本次已完成範圍。設計記錄於 `docs/runbooks/agentic-monorepo-detection.md`，每次 run 會送出 `wizard: agentic detection` event 觀察 rollout 結果。

GitHub Actions 可用以下 Wizard CI commands：

```text
/wizard-ci all
/wizard-ci basic-integration
/wizard-ci mcp-analytics
/wizard-ci revenue
/wizard-ci basic-integration/android
/wizard-ci basic-integration/angular
/wizard-ci basic-integration/astro
/wizard-ci basic-integration/django
/wizard-ci basic-integration/fastapi
/wizard-ci basic-integration/flask
/wizard-ci basic-integration/javascript-node
/wizard-ci basic-integration/javascript-web
/wizard-ci basic-integration/laravel
/wizard-ci basic-integration/next-js
/wizard-ci basic-integration/nuxt
/wizard-ci basic-integration/python
/wizard-ci basic-integration/rails
/wizard-ci basic-integration/react-native
/wizard-ci basic-integration/react-router
/wizard-ci basic-integration/sveltekit
/wizard-ci basic-integration/swift
/wizard-ci basic-integration/tanstack-router
/wizard-ci basic-integration/tanstack-start
/wizard-ci basic-integration/vue
/wizard-ci mcp-analytics/custom-dispatcher
/wizard-ci mcp-analytics/typescript-sdk
/wizard-ci revenue/stripe
```

**Queue 與 tool orchestration** `queue-tools.ts`（commit `ba91f26c418f332f1ede8b6e80fd9fa14bbb22e8`）位於 PostHog/wizard 的原始檔案，共 `479 lines`、`444 loc`、`16.4 KB`。它在存在 queue 時向既有 `wizard-tools` server 註冊 `enqueue_task`、`complete_task` 與 `read_handoffs`，讓 orchestrator agent 與 task agents 以 structured handoff 協作。 

![](https://pub-75d4fe1e4e80421b9ecb1245a7ae0d1a.r2.dev/curated/2944653337e811b4.jpg)
> Slack 對話截圖記錄了一則關於 CI 執行失敗與非確定性問題的討論，其中包含了發文者 Vincent 與 edwin 的訊息，並引用了一筆 GitHub Pull Request 連結 `#1870 [CI] (9ab9173) nuxt/movies-nuxt-3-6`，附帶 PostHog/wizard-workbench 儲存庫標籤及 CI/CD 標籤。

`validTypes` 會拒絕未知 task type；`sinkTypes` 必須最後執行，且 transitively depend on queue 中所有其他 task；`runnerSeededTypes` 會 deferred to the end of the drain，可能要求 user input，使下游等待。queue 上限是 `MAX_QUEUE_TASKS = 30`，正常 flow 約 `9 tasks`，這只是防止 runaway growth 的 backstop。dedup key 由 task `type` 與穩定排序的 `inputs` 組成：

```typescript
const MAX_QUEUE_TASKS = 30
function dedupKey(type: string, inputs: Record<string, unknown>): string {
  return ` ${type} :: ${stableStringify(inputs)} `
}
```

`complete_task` 要求每個 task 結束時「Always call this exactly once」，status 可為 `done`、`failed` 或 `not needed`；不適用的 task 必須使用 `not needed` 並說明原因。guard 失敗會呼叫 `analytics.wizardCapture('orchestrator guard tripped', { guard, type })`，而不是讓流程無限重試。

**工具呼叫也要測試完整輸出** GPT-5.6 的 [Model guidance｜OpenAI API](https://developers.openai.com/api/docs/guides/latest-model?model=gpt-5.6) 建議，使用 Programmatic Tool Calling（PTC）時，prompt 必須明確定義 bounded stage、可用 tools、exact output schema、evidence、concurrency、retry 與 stopping 限制；若無法預先判斷回傳資料 shape，應改用 direct tool calling。PTC 與 direct tool calling 並存時，要定義清楚 handoff，避免切換路徑或重做已完成工作。

評估不能只看 program 是否成功，還要分開檢查 `program_output` 與最後的 assistant message，確認 required field、citation、evidence 與 caveat 都存在；並在相同 representative tasks 上比較 total tokens、latency、cost、calls、turns 與 retries，只有通過既有 evals 才算改善。

**失敗案例揭露文件腐化風險** PostHog 的 [PR #75873](https://github.com/PostHog/posthog/pull/75873) 顯示，Trunk 暫停後，`AGENTS.md` 仍要求所有合併走 Trunk merge queue 並禁止 `gh pr merge`。Twixes 實際浪費約 45 分鐘，最後以 repository 唯一允許的 squash merge 成功：

```bash
gh pr merge <number> --squash
gh api repos/PostHog/posthog/rulesets
hogli ci:preflight --fix
```

當時「Trunk merge」ruleset 為 `disabled`，`repository.mergeQueue(branch: "master")` 回傳 `null`；但 approving review、code owner review、required status checks、signed commits 等 branch protection 仍有效。2026 年 8 月 2 日，gantoine 又提到 PR #76456「chore(ci): revert Trunk pause docs and gate desktop Trunk uploads」，因此 7 月 31 日的文件修正不能視為永久結論。

**PR 與 skill 的治理** PostHog 的 PR template 要求說明可獨立閱讀，清楚分開 `Problem`、`Changes`、測試與限制；Agent 不得聲稱未執行的 manual testing。Autonomy 只能填 `Human-driven (agent-assisted)` 或 `Fully autonomous`，並列出 Agent、tool、session link、所有使用的 public 或 repo-provided `skills`，不可貼 user prompt 或敏感資料。Agent-authored PR 必須有人 review，不得 self-merge 或 auto-approve。

Template 也禁止未經確認就上傳資料。`hogli pr:upload-image <file>` 會將圖片上傳至公開的 `PostHog/pr-assets` repository，只能上傳不含 customer data、secrets 或 internal info 的檔案；這類外部上傳指令屬高風險，需人工核對內容後再執行。建立 PR 的指令為：

```bash
gh pr create --draft
gh pr create --body-file -
```

流程或 topology 變更還必須提供獨立的 before／after Mermaid `flowchart` blocks；建立 PR 預設使用 draft，修正 CI 並執行 affected tests 後才標記 ready for review。

**持續修補與驗證** context-mill PR [#272](https://github.com/PostHog/context-mill/pull/272) 由 `posthog[bot]` 於 2026 年 7 月 22 日提出，根據 wizard remark telemetry 

![](https://pub-75d4fe1e4e80421b9ecb1245a7ae0d1a.r2.dev/curated/a49e69bad19005e2.jpg)
> PostHog 中記載 remark 字串與多筆任務執行備忘文字的表格介面

 聚類修正不存在的 `products-enable` MCP tool、deprecated aliases、notebook JSON escaping 與 Python constructor 規則矛盾；測試結果為 `137 tests pass`、full build succeeds：

```bash
npm test
npm run build
```

但 `init-not-duplicated` check ID mismatch 未納入範圍，不能宣稱已解決。

另一個已合併的 PR [#81018](https://github.com/PostHog/posthog/pull/81018) 修正 `.agents/skills/clickhouse-migrations/SKILL.md` 虛構 `NodeRole.COORDINATOR` 的問題。該錯誤曾讓四個 PR 的 backend、Dagster 與 code-quality jobs 全部變紅；修正後加入 satellite cluster 判斷，並要求 managed role 同步更新 `posthog/clickhouse/hcl/`。作者以實際 migration 驗證後執行：

```bash
hogli lint:skills
gen-golden.sh
gen-sql.sh
```

`hogli lint:skills` 通過全部 `143 skills`，另有 `3` 個既存 advisory warnings。這些案例共同顯示，context 的品質不能只靠文字審查：它必須與 repository、CI、GitHub settings、實際執行結果及 Agent 回饋互相驗證。 

![](https://pub-75d4fe1e4e80421b9ecb1245a7ae0d1a.r2.dev/curated/2e00eeaef9bea030.png)
> PostHog 的亮橘色長條形按鈕，中央以白色粗體字寫著「build mode」，左側帶有雙快轉箭頭圖示，背景為淺藍色網格線，右下角為翻起頁角與 PostHog logo。

## 標籤

功能更新, 教學資源, Claude Code
