每个上线的LLM功能最后都会撞上同一堵墙。一开始它只是Python文件里的一行提示词字符串,然后长出系统消息、三个示例、一段对付JSON偶尔被markdown包住的格式化补丁,最后留下一行注释:# 不要动——2026-08-14调过。没人能测它,没人能diff它,模型厂商一发更新,它就悄悄退化。
DSPy——来自斯坦福NLP的框架,MIT许可,当前版本3.4.0——从根上拆这堵墙。你不再写提示词,而是写程序:用带类型的签名声明输入输出,用模块组合推理策略,再用优化器针对你定义的指标把整个东西编译一遍,就像编译器针对目标架构优化代码。提示词变成了构建产物,不再是源代码。
一条端到端的回路
本教程要构建的正是这样一条回路:一个客服工单分诊程序,预测紧急程度和负责团队。用签名定义它,用dspy.Evaluate测出真实基线,用BootstrapFewShot编译它,观察准确率从55.6%走到88.9%,最后把优化后的程序存成文件发出去。下面的一切都实际执行过——包括失败模式——整条流水线免费、离线,跑在一个确定性的沙盒模型上,零API开销就能复现每一个数字。
配图:你将构建的DSPy回路——写程序、测量、编译。
需要准备什么
Python 3.10+ 和 pip。CPU就够,这里不需要GPU。
DSPy 3.4.0:pip install dspy,然后用 python -c "import dspy; print(dspy.__version__)" 确认,应该输出3.4.0。
主流水线不需要API密钥:它跑在SandboxLM上,一个下文给出的小型确定性模拟器,用来顶替真实模型。它从提示词里读取少样本示例,读不到就退回一个弱关键词启发式——故意做得平庸,就像一个真实的零样本基线。
可选,想接真实模型时:任意模型API密钥(OpenAI、Anthropic或本地Ollama服务),一行代码就能在第一步把沙盒换成真实模型。
第一步:安装与配置
先安装并验证版本:
pip install dspy
python -c "import dspy; print(dspy.__version__)" # 3.4.0
每个DSPy程序都从告诉框架用哪个语言模型开始,通过dspy.configure完成。接真实服务商时,dspy.LM接受任何LiteLLM风格的模型字符串,密钥从环境变量读取:
import dspy
# 真实模型路线(需要环境里有OPENAI_API_KEY):
# dspy.configure(lm=dspy.LM("openai/gpt-4o-mini"))
# 本地路线(需要Ollama在运行):dspy.LM("ollama/llama3.1", api_base="http://localhost:11434")
本教程里配置的是沙盒。把下面这段存成sandbox_lm.py:它继承DSPy的DummyLM(来自dspy.utils.dummies),并重写_use_example钩子——这是DSPy 3.4的dummy引擎实际调用的扩展点。给定一个查询,它会找出输入内容词重叠最多的那个少样本示例,复制该示例的输出;没有重叠示例时,退回弱启发式:
"""Deterministic stand-in for a real LM. Reads few-shot demos from its prompt."""
import re
import dspy
from dspy.utils.dummies import DummyLM
HDR = re.compile(r"\[\[ ## (\w+) ## \]\]")
STOP = set("""a an the and or of to in on for with is are was were be been
it its this that these those i my we you your he she they them his her our
at as by from have has had do does did will would can could should there
their what when where which who how why not no yes if then than so such
very just about into over after before between me us him her them""".split())
def _split_fields(text):
parts = HDR.split(text)
return [(parts[i], parts[i + 1].strip())
for i in range(1, len(parts) - 1, 2)]
def _tokens(text):
return [t for t in re.findall(r"[a-z0-9]+", text.lower()) if t not in STOP]
def _heuristic(ticket):
t = ticket.lower()
if any(w in t for w in ["invoice", "charge", "charged", "billing",
"refund", "payment", "subscription", "receipt"]):
team = "billing"
elif any(w in t for w in ["package", "tracking", "delivery", "shipment",
"arrived", "parcel", "shipping"]):
team = "shipping"
else:
team = "technical"
urgency = ("high" if any(w in t for w in ["urgent", "immediately", "asap",
"locked", "down", "breach", "twice", "double", "angry",
"cancel", "fraud", "lost"]) else "low")
return {"urgency": urgency, "team": team}
class SandboxLM(DummyLM):
def __init__(self):
super().__init__([], follow_examples=True)
def _use_example(self, messages):
users = [m["content"] for m in messages if m["role"] == "user"]
assistants = [m["content"] for m in messages if m["role"] == "assistant"]
input_names = {n for n, _ in _split_fields(users[0])}
last = users[-1]
out_names = [n for n, _ in _split_fields(last) if n not in input_names]
query = " ".join(v for n, v in _split_fields(last) if n in input_names)
qtok = set(_tokens(query))
demos, ai = [], 0
for u in users[:-1]:
a = assistants[ai]; ai += 1
outs = {n: v for n, v in _split_fields(a) if n in out_names}
if outs:
demos.append((" ".join(v for n, v in _split_fields(u)
if n in input_names), outs))
best, best_score = None, 0
for dtext, outs in demos:
score = len(qtok & set(_tokens(dtext)))
if score > best_score:
best, best_score = outs, score
if best is not None and best_score >= 2:
values = {n: best.get(n, "") for n in out_names}
else:
h = _heuristic(query)
values = {n: h.get(n, "Reading the ticket and the closest examples.")
if n in h or n in ("reasoning", "rationale") else ""
for n in out_names}
return self._format_answer_fields(values)
然后配置它:
from sandbox_lm import SandboxLM
dspy.configure(lm=SandboxLM())
有一点要说清楚:这个模拟器不是LLM。它只模拟真实模型的一个属性——从上下文示例中做少样本学习——好让优化回路的行为和面对真实模型时一致,下文每一次测量都(原文在此处截断)
热门跟贴