Signals
从三个 signal 开始,再参考另外九个:解释失败、识别说谎的 200 响应、agent 回合、审计操作和不稳定任务。

这里的每个 signal 都是在仅凭字段无法回答时提出的问题。将 when 改成你的路径,它们就属于你的应用。

从这三个开始

一个文件、三列,运行一周后你就会拥有以前没有的数据。

server/signals.ts
import { defineSignal } from '@evlog/signals'

const errorName = (e: Record<string, unknown>) => (e.error as { name?: string } | undefined)?.name ?? 'none'

/** Every 4xx/5xx. Splits the error budget between the client, the code and the dependencies. */
export const fault = defineSignal({
  name: 'fault',
  when: e => (e.status ?? 0) >= 400,
  ask: 'Who is responsible for this failure?',
  choice: {
    client: 'Bad input, expired session, missing permission, client mistake',
    app: 'A bug, a misconfiguration or a validation error in our own code',
    upstream: 'A third-party dependency failed, timed out or rate-limited us',
  },
  cacheKey: e => e.error ? `${e.method} ${e.path} ${e.status} ${errorName(e)}` : undefined,
})

/** Every 5xx. Tells a retry policy which failures are worth a second attempt. */
export const retryable = defineSignal({
  name: 'retryable',
  when: e => (e.status ?? 0) >= 500,
  ask: 'Would the same request most likely succeed if retried in a few seconds?',
  criteria: { true: 'Timeout, connection reset, rate limit, transient upstream error', false: 'Bug, bad data, missing resource' },
  cacheKey: e => e.error ? `${e.path} ${errorName(e)}` : undefined,
})

/** Successful checkouts. The 200 that sampling deletes and nobody notices. */
export const silentFailure = defineSignal({
  name: 'silent-failure',
  when: e => e.status === 200 && e.path === '/api/checkout',
  ask: 'The request returned 200, but the customer did not get what they came for',
  criteria: { true: 'No order or confirmation, a fallback path, an empty or partial result', false: 'Order created and confirmed' },
  keep: v => v.value && v.confidence > 0.8,
})

每个 signal 提供的结果:

  • 一周内执行 GROUP BY signals.fault.value,即可知道下一个 sprint 应该投入验证消息、自身 bug,还是依赖项的 SLA。在错误名称上设置 cacheKey,可以让一次故障只产生一次调用,而不是每个请求一次。
  • 502 上的 signals.retryable.value = true,就是重试策略与猜测之间的区别。
  • silent-failure 会提升返回 200 却没有订单的结账请求。将 when 指向你自己的资金路径。

全部十二个

Signal钩子问题规则为什么无法回答
faultenrich、缓存责任属于用户、我们还是上游Stripe 超时和空值解引用都会得到 500
severityenrich是噪声、需要关注还是需要通知值班人员紧急程度取决于失败内容和受影响的人
retryableenrich、缓存重试会成功吗临时错误和永久错误可能共享状态码
slow-causeenrich数据库、上游还是计算读取时间分布需要人工判断
silent-failurekeep返回了 200,但客户空手而归吗采样只能保留谓词能够命名的内容
webhook-ignoredkeep已确认但没有执行吗provider 只能看到 200
validation-bugkeep被拒绝的输入其实有效吗看起来与无效输入相同
turn-resolvedenrichagent 完成了要求吗抽样的 LLM-as-judge 会漏掉其余回合
turn-loopingkeep是否重复调用工具却没有进展循环会隐藏在成功回合中
turn-off-scriptkeep是否执行了没人要求的操作同上
audit-reviewkeep此操作是否值得人工审核风险来自字段组合,而不是某个字段
job-flakyenrich是否只是因为重试才成功重试会隐藏原因

其他适合的问题

first-seen(此错误对该路由来说是新的吗)、user-impact(无影响、降级、阻塞)、owner(哪个团队应该处理)、known-error(根据你自己的 why 和 fix 判断属于哪个目录项)、支持对话中的 frustration、会话事件中的 did-succeed。

不适合的问题

  • 机密或 PII 检测。 把值发送给模型,再询问它是否为机密本身就是泄露。请使用脱敏。
  • 通知值班人员。 永远不要根据概率触发通知。severity 负责排列优先阅读的内容;根据 status 和 durationMs 设置的告警才决定叫醒谁。
  • 将对抗性内容作为单一判定。 能够被提问的模型也能够被诱导。
  • 数值问题。 决策模型返回选项分布,而不是数字。请在处理程序中计算并记录数字。