# Shopify 以 Qwen3.5-0.8B 將 Buyer Profile 產能提高 36 倍；Flow 另用 Qwen3-32B 與 Tangle 建立 production 回饋迴圈

> 📖 本站完整內容索引（documentation index）：[llms.txt](/llms.txt)

> 原作者：tobi lutke (@tobi) · 策展與摘要：EasyVibeCoding · 平台：X (Twitter) · 熱度：🔥🔥🔥🔥 · 日期：2026-09-02

> 原始來源：https://x.com/tobi/status/2094904650709234015

## 證據與延伸閱讀

- [Shopify 以 Qwen3.5-0.8B 將 Buyer Profile 產能提高 36 倍；Flow 另用 Qwen3-32B 與 Tangle 建立 production 回饋迴圈。](https://tangleml.com/) — 官方文件 · 最後核對：2026-09-02 · 支持主張：新增 Tangle 作為可重現、具容器化元件與 content-based caching 的資料／訓練／評估實驗圖，並區分兩個不可混淆的案例：Buyer Profile slide 報告 GPT-5.6-sol xhigh teacher 83.0 到 Qwen3.5-0.8B student 84.6、prompt 9.1K 到 1.1K、以及 2M 到 72M profiles/day on 100 H100 GPUs；Flow Qwen3-32B case study 則以 300 hand-crafted examples、syntax +22、semantic +13、1% traffic activation -35%、2.2-times latency、68% lower cost 與兩週 feedback flywheel 展示 production loop。TangleML/tangle 公開 backend 且採 Apache License 2.0；Buyer Profile 的 judge prompt、test split、confidence inte…
- [Shopify Tangle with Buyer Profile Qwen3.5-0.8B and Flow Qwen3-32B — github.com](https://github.com/TangleML/tangle) — 官方 Repository · 最後核對：2026-09-02 · 支持主張：新增 Tangle 作為可重現、具容器化元件與 content-based caching 的資料／訓練／評估實驗圖，並區分兩個不可混淆的案例：Buyer Profile slide 報告 GPT-5.6-sol xhigh teacher 83.0 到 Qwen3.5-0.8B student 84.6、prompt 9.1K 到 1.1K、以及 2M 到 72M profiles/day on 100 H100 GPUs；Flow Qwen3-32B case study 則以 300 hand-crafted examples、syntax +22、semantic +13、1% traffic activation -35%、2.2-times latency、68% lower cost 與兩週 feedback flywheel 展示 production loop。TangleML/tangle 公開 backend 且採 Apache License 2.0；Buyer Profile 的 judge prompt、test split、confidence inte…
- [Shopify Tangle with Buyer Profile Qwen3.5-0.8B and Flow Qwen3-32B — shopify.engineering](https://shopify.engineering/fine-tuning-agent-shopify-flow) — 官方文件 · 最後核對：2026-09-02 · 支持主張：新增 Tangle 作為可重現、具容器化元件與 content-based caching 的資料／訓練／評估實驗圖，並區分兩個不可混淆的案例：Buyer Profile slide 報告 GPT-5.6-sol xhigh teacher 83.0 到 Qwen3.5-0.8B student 84.6、prompt 9.1K 到 1.1K、以及 2M 到 72M profiles/day on 100 H100 GPUs；Flow Qwen3-32B case study 則以 300 hand-crafted examples、syntax +22、semantic +13、1% traffic activation -35%、2.2-times latency、68% lower cost 與兩週 feedback flywheel 展示 production loop。TangleML/tangle 公開 backend 且採 Apache License 2.0；Buyer Profile 的 judge prompt、test split、confidence inte…
- [Shopify Tangle with Buyer Profile Qwen3.5-0.8B and Flow Qwen3-32B — shopify.engineering](https://shopify.engineering/tangle) — 官方文件 · 最後核對：2026-09-02 · 支持主張：新增 Tangle 作為可重現、具容器化元件與 content-based caching 的資料／訓練／評估實驗圖，並區分兩個不可混淆的案例：Buyer Profile slide 報告 GPT-5.6-sol xhigh teacher 83.0 到 Qwen3.5-0.8B student 84.6、prompt 9.1K 到 1.1K、以及 2M 到 72M profiles/day on 100 H100 GPUs；Flow Qwen3-32B case study 則以 300 hand-crafted examples、syntax +22、semantic +13、1% traffic activation -35%、2.2-times latency、68% lower cost 與兩週 feedback flywheel 展示 production loop。TangleML/tangle 公開 backend 且採 Apache License 2.0；Buyer Profile 的 judge prompt、test split、confidence inte…
- [Tobi Lutke Buyer Profile slide 顯示 Qwen3.5-0.8B student 拿到 84.6 分](https://x.com/tobi/status/2094904650709234015) — 一手來源
- [模型分數由 7 月 23 日的 75.3 提升至 7 月 30 日的 84.6](https://cdn.shopify.com/s/files/1/0779/4361/files/LLM_Judge_score_over_each_month.png?v=1776860795) — 一手來源

## 中文摘要

Shopify 以 Qwen3.5-0.8B 將 Buyer Profile 產能提高 36 倍；Flow 另用 Qwen3-32B 與 Tangle 建立 production 回饋迴圈。

**Buyer Profile 結果** Tobi Lutke 分享的 Buyer Profile slide 顯示，Qwen3.5-0.8B student 在特定任務拿到 84.6，較 GPT-5.6-sol xhigh teacher 的 83.0 高 1.6 個百分點；原本 9.1K-token 的 system prompt 壓到 1.1K tokens，約剩八分之一。在 100 張 H100 上，處理量從每日 2M profiles 成長為 72M，成長為 36 倍。過程中，模型分數由 7 月 23 日的 75.3（29K samples），提升至 7 月 27 日的 78.1（42K samples），再到 7 月 30 日的 84.6（54K samples）；Qwen3.5-2B production baseline 則為 77.2。

![](https://pub-75d4fe1e4e80421b9ecb1245a7ae0d1a.r2.dev/curated/cbb1496474bdcf31.jpg)
> Qwen3.5-0.8B 在 54k 樣本微調後 Judge score 達到 84.6，超越 GPT-5.6-sol (xhigh) Teacher 的 83.0 與前代生產模型 Qwen3.5-2B 的 77.2，同時 system prompt 壓縮至 1.1K tokens，每日產出 Profile 提升至 72M。

**Flow production loop** Shopify 另一篇 Flow case study 使用的是 Qwen3-32B，不是 Buyer Profile 的 Qwen3.5-0.8B；兩案使用的資料、評估器、流量與指標都不同，不能把分數或效能數字互相比高低。Flow 將商家的自然語言轉成會呼叫工具的自動化流程。團隊以 300 個人工設計樣本評估語意正確性、語法與延遲；把目標格式從巢狀 JSON DSL 改為 Python 後，語法正確性提高 22 個百分點、語意正確性提高 13 個百分點。模型訓練需在兩個 H200 節點上執行 12 小時。

![](https://pub-75d4fe1e4e80421b9ecb1245a7ae0d1a.r2.dev/curated/9dfbc62eec47e6b3.png)
> Qwen3-32B fine-tuned 在 LLM Judge Overall Pass Rate 指標上自 Jan 的 22.9% 升至 April 的 56.1%，並在 April 領先 Closed-Source Frontier Model 的 51.0%。

**真實流量驗證** Flow 首次只開放 1% 流量時，即使離線基準測試與原方案相當，實際工作流啟用率仍低 35%，顯示離線測試可能漏掉生產環境的特定失敗情境。Shopify 以數百筆人工標註對話校準 LLM 評分器，接著每週匯入 production 對話並評分，把高品質樣本送進訓練、隔離低品質樣本、找出缺口、重新訓練再部署。Shopify 表示這套 flywheel 在兩週內補上 offline 與 production 的品質落差；目前上線的 32B Flow agent 相較它取代的 frontier model 快 2.2 倍、成本低 68%。官方沒有公開兩週的起訖日期，也沒有提供成本前後值或計算式，因此這些仍是第一方 production 結果。

**Tangle 基礎設施** Tangle 是 Shopify ML 與軟體工程師打造的平台無關、可重現的 AI pipeline 系統，不是 learning algorithm。其視覺編輯器把可重用元件組成 pipeline，元件可在容器內使用 Python、Java、shell、Ruby、C++ 或 JavaScript/TypeScript；content-based caching 會重用未受變更影響的 upstream outputs，讓原本約 10 小時的流程在只改動單一元件時縮短至約 20 分鐘。公開的 [TangleML/tangle backend 程式庫](https://github.com/TangleML/tangle)採 Apache License 2.0。

![](https://pub-75d4fe1e4e80421b9ecb1245a7ae0d1a.r2.dev/curated/14538ba31e602bdd.png)
> Tangle Pipeline 的運作流程圖，大框內包含 Data Pipeline 連至 Training（2x H200 / 12 hours）、Evaluation LLM Judge 與 Deploy CentML，右側外接 HuggingFace Datasets & Models 以及下方 CometML Experiment Tracking。

**重現邊界** Tangle 提供的是資料、訓練與評估的迭代編排；若沒有評估器、資料分流政策、隔離流程、重新訓練與部署門檻，模型不會自行改善。Buyer Profile 尚未提供評分提示詞、測試切分、信賴區間、訓練成本、teacher 輸出、資料集、可重現的公開 checkpoint 或完整做法，因此 84.6 分仍是第一方報告，無法只靠目前公開的資料獨立重現。

## 標籤

功能更新, Shopify, Qwen3-32B
