本文参考了 Microsoft Community Hub 上 Robin Lester 的文章 Using Jev with Agents in Microsoft Foundry for Model Evaluation,下面是我整理的理解和步骤。

Jev 是什么

Jev 是 TypeSafe AI 推出的新模型,2026 年 9 月 15 日开放了抢先体验。TypeSafe 把它称为自家第一个“System One”模型,它不写长篇文字,只负责又快又结构化地做判断。

用法很直接:把一段状态或上下文交给它,再说明你想让它做什么判断,它会返回一个带类型的概率化结果。结果有三种形式:

  • Choice:从给定的选项里选一个
  • Score:在一个有序的刻度上打分
  • Yes/No 概率:判断某个条件成立的概率

普通 LLM 吐出的是字符串,Jev 放弃了这一点,换来类型安全的结构化数值。TypeSafe 自己的说法是:输入非结构化的状态,输出带类型的概率化决策。

为什么要用 Jev

当应用需要的是一次判断,而不是一段回答时,Jev 最能派上用场。比如:

  • 按明确定义的标准做一致的评估
  • 输出结构化结果,应用可以直接消费
  • 分类、评估、路由和决策支持
  • 在 Agent 工作流里,先判断再决定要不要执行动作
  • 做护栏,例如判断 Agent 该不该发起某次工具调用
  • 做模型路由,由一次判断决定请求交给哪个模型或流程

每个决策还可以带上置信度或概率。有了这个数,应用就能设阈值:高于阈值自动往下走,低于阈值转人工复核。

Jev 与可解释性

对于 Model-as-a-Service(MaaS)和托管的 LLM 方案,Jev 提供了另一种解释思路。

SHAP 这类传统的可解释性技术给出的是特征归因,解释模型为什么做出某个预测。但模型只通过托管 API 访问、拿不到内部结构时,这些技术就很难用上。

Jev 不能取代 SHAP,给不出同类的特征归因。它的位置是一个独立的判断层:按明确的标准去评估 LLM 的输入、输出或打算执行的动作,再返回带置信度的结构化决策。流程大概是这样:

LLM 生成答案 → Jev 评估答案 → 应用决定下一步

对那些看不透内部的 MaaS 模型来说,这等于在外面套了一层可度量的评估。

换成另一个 LLM 行不行

行。你完全可以让普通 LLM 对输入做分类、打分或评估,并要求它输出 JSON。

区别在于 Jev 是专门为这类任务造的,不是通用的文本生成模型。TypeSafe 称 Jev 并行产出结果,而不是逐个 token 顺序生成,托管服务目前标称的典型决策延迟约为 70 到 500 毫秒。

价格

截至 2026 年 9 月,TypeSafe 文档里列出的 Jev 1.13(jev-1.13.0)价格是每百万输入 token 0.042 美元,输出 token 免费。限速是每秒 250,000 个 token、每分钟 1,200 次请求,不过 TypeSafe 说明这些限制可能调整。

Jev 产出的是决策而不是文本,所以你付费的对象是它要评估的上下文,而不是生成的内容。

价格可能变动,以上数字仅在原文发布时有效。

在 Microsoft Foundry 里使用 Jev

我觉得比较实用的一种搭法,是在 Microsoft Foundry 里把 Jev 当成生成式模型旁边的专用判断组件:

用户/应用 → Foundry 模型或 Agent → Jev 判断 → 确定性逻辑 → 动作/响应

举个例子:Foundry 上托管的模型先生成一个候选答案,Jev 按预设标准评估这个答案,应用再根据评估结果决定流程能不能继续。

具体配置步骤如下。

申请 Jev 并创建 API Key

先向 TypeSafe 申请 Jev 的预览资格。有了账号以后,需要创建一个 API Key。

在菜单里进入 API Keys,新建一个 Key,并把它复制下来,后面要在 Foundry 里用。

TypeSafe 控制台的 API Keys 页面

你可能还会收到一些免费额度。Jev 的单价很低,拿来做实验绰绰有余。

在 Foundry 里创建 Agent 并添加 OpenAPI 工具

接下来到 Foundry 里为 Jev 新建一个 Agent。在 Agent 里添加一个 Tool,选择 OpenAPI tool,如下图所示。

在 Foundry 的 Select a tool 窗口中选择 OpenAPI tool

然后按下图填写对话框。

Update OpenAPI tool 对话框的填写示例

这里需要新建一个连接。

新建连接时填写的 Credential

Key 的格式是 Bearer apikey_<你的 Key 的其余部分>。

Schema 部分填入下面这段:

  1
  2
  3
  4
  5
  6
  7
  8
  9
 10
 11
 12
 13
 14
 15
 16
 17
 18
 19
 20
 21
 22
 23
 24
 25
 26
 27
 28
 29
 30
 31
 32
 33
 34
 35
 36
 37
 38
 39
 40
 41
 42
 43
 44
 45
 46
 47
 48
 49
 50
 51
 52
 53
 54
 55
 56
 57
 58
 59
 60
 61
 62
 63
 64
 65
 66
 67
 68
 69
 70
 71
 72
 73
 74
 75
 76
 77
 78
 79
 80
 81
 82
 83
 84
 85
 86
 87
 88
 89
 90
 91
 92
 93
 94
 95
 96
 97
 98
 99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
{
  "openapi": "3.0.3",
  "info": {
    "title": "TypeSafe Jev System One",
    "version": "1.1.0",
    "description": "Evaluate content using TypeSafe Jev System One."
  },
  "servers": [
    {
      "url": "https://api.typesafe.ai"
    }
  ],
  "paths": {
    "/v1/systemone": {
      "post": {
        "operationId": "evaluate_with_jev",
        "summary": "Evaluate content using Jev",
        "description": "Evaluate state against typed questions using TypeSafe Jev.",
        "security": [
          {
            "bearerAuth": []
          }
        ],
        "requestBody": {
          "required": true,
          "content": {
            "application/json": {
              "schema": {
                "type": "object",
                "additionalProperties": false,
                "required": [
                  "state",
                  "model",
                  "questions"
                ],
                "properties": {
                  "state": {
                    "type": "string",
                    "description": "The original text or content for Jev to evaluate."
                  },
                  "model": {
                    "type": "string",
                    "enum": [
                      "jev-latest"
                    ],
                    "default": "jev-latest",
                    "description": "The TypeSafe Jev model. Use jev-latest."
                  },
                  "questions": {
                    "type": "object",
                    "description": "A map of named questions for Jev to answer. Each question must be a Noul, Choice, or Score question.",
                    "additionalProperties": {
                      "oneOf": [
                        {
                          "$ref": "#/components/schemas/NoulQuestion"
                        },
                        {
                          "$ref": "#/components/schemas/ChoiceQuestion"
                        },
                        {
                          "$ref": "#/components/schemas/ScoreQuestion"
                        }
                      ]
                    }
                  }
                }
              }
            }
          }
        },
        "responses": {
          "200": {
            "description": "Successful Jev evaluation",
            "content": {
              "application/json": {
                "schema": {
                  "type": "object",
                  "additionalProperties": true
                }
              }
            }
          },
          "400": {
            "description": "Invalid request"
          },
          "401": {
            "description": "Authentication failed"
          },
          "422": {
            "description": "Request validation failed"
          }
        }
      }
    }
  },
  "components": {
    "schemas": {
      "NoulQuestion": {
        "type": "object",
        "additionalProperties": false,
        "required": [
          "type",
          "instructions"
        ],
        "properties": {
          "type": {
            "type": "string",
            "enum": [
              "noul"
            ]
          },
          "instructions": {
            "type": "string",
            "description": "The yes/no question for Jev to evaluate."
          },
          "criteria": {
            "type": "object",
            "additionalProperties": false,
            "properties": {
              "true": {
                "type": "string",
                "description": "Description of what a yes or value near 1 means."
              },
              "false": {
                "type": "string",
                "description": "Description of what a no or value near 0 means."
              }
            }
          }
        }
      },
      "ChoiceQuestion": {
        "type": "object",
        "additionalProperties": false,
        "required": [
          "type",
          "instructions",
          "criteria"
        ],
        "properties": {
          "type": {
            "type": "string",
            "enum": [
              "choice"
            ]
          },
          "instructions": {
            "type": "string",
            "description": "The question for Jev to decide between the supplied choices."
          },
          "criteria": {
            "type": "object",
            "description": "A map where each property name is a possible choice and its value describes that choice.",
            "minProperties": 2,
            "additionalProperties": {
              "type": "string"
            }
          }
        }
      },
      "ScoreQuestion": {
        "type": "object",
        "additionalProperties": false,
        "required": [
          "type",
          "instructions",
          "criteria"
        ],
        "properties": {
          "type": {
            "type": "string",
            "enum": [
              "score"
            ]
          },
          "instructions": {
            "type": "string",
            "description": "The attribute or question Jev should score."
          },
          "criteria": {
            "type": "array",
            "description": "Ordered scoring levels from lowest to highest.",
            "minItems": 2,
            "items": {
              "type": "string"
            }
          }
        }
      }
    },
    "securitySchemes": {
      "bearerAuth": {
        "type": "apiKey",
        "name": "Authorization",
        "in": "header"
      }
    }
  },
  "security": [
    {
      "bearerAuth": []
    }
  ]
}

给 Agent 写指令

这个例子里我们用 Jev 来评估客户投诉,所以 Agent 的指令如下:

 1
 2
 3
 4
 5
 6
 7
 8
 9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
You are a customer-support triage agent.

For every customer-support message you MUST call the evaluate_with_jev tool.

IMPORTANT:
When calling evaluate_with_jev, construct Jev questions using ONLY the following property names:

For a Noul question:
{
  "type": "noul",
  "instructions": "...",
  "criteria": {
    "true": "...",
    "false": "..."
  }
}

For a Choice question:
{
  "type": "choice",
  "instructions": "...",
  "criteria": {
    "option1": "...",
    "option2": "..."
  }
}

For a Score question:
{
  "type": "score",
  "instructions": "...",
  "criteria": [
    "lowest level",
    "middle level",
    "highest level"
  ]
}

NEVER use these properties:
- question
- prompt
- options
- ordered_criteria

Always use:
- instructions
- criteria

For customer support tickets, ask Jev these three questions:

billing:
type = noul
instructions = "Does this ticket concern billing, payments, charges, invoices or refunds?"
criteria.true = "The issue concerns billing or a financial transaction."
criteria.false = "The issue does not concern billing or a financial transaction."

tone:
type = choice
instructions = "What is the customer's tone?"
criteria:
calm = "Neutral, polite or informational."
frustrated = "Dissatisfied, impatient or inconvenienced."
angry = "Strongly confrontational or extremely dissatisfied."

urgency:
type = score
instructions = "How urgently does this ticket require attention?"
criteria in this exact order:
1. "Can wait"
2. "Handle this week"
3. "Handle today"

Always set model to "jev-latest".
Pass the customer's original message as state.
Do not classify the ticket yourself.
Use the values returned by Jev.
Always show the Jev output at the end of the response

试一条投诉

现在给它一条投诉,看看 Jev 怎么判断:

1
The room had no water in and one of the towels was dirty

Agent 的回复如下:

 1
 2
 3
 4
 5
 6
 7
 8
 9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
I’m sorry about the issues you experienced—no running water and a dirty towel are unacceptable. I’ll get this addressed right away.

Immediate help I can arrange now:
- Send maintenance to restore water.
- Have housekeeping deliver fresh towels and re-service the bathroom.
- Provide bottled water immediately.
- If water can’t be restored quickly, arrange a room move.

Please share:
- Your reservation name and room number
- Best contact number
- Whether there’s no water at all or only hot/cold is affected
- Whether you prefer immediate service or a room change

If you’ve already left, let me know the date/time this happened so we can investigate and follow up.

Jev output:
{
  "model": "jev-1.13.0",
  "answers": {
    "billing": {
      "type": "noul",
      "noul": 0.02
    },
    "tone": {
      "type": "choice",
      "choice": "frustrated",
      "confidence": 0.6,
      "probabilities": {
        "frustrated": 0.74,
        "calm": 0.26,
        "angry": 0.0
      }
    },
    "urgency": {
      "type": "score",
      "score": 1.81,
      "confidence": 0.72,
      "legend": {
        "0": "Can wait",
        "1": "Handle this week",
        "2": "Handle today"
      },
      "probabilities": {
        "0": 0.03,
        "1": 0.12,
        "2": 0.85
      }
    }
  },
  "usage": {
    "input_tokens": 466,
    "output_tokens": 75
  }
}

这样 Jev 的决策和背后的概率就都看得到了。这条投诉几乎不可能和账单有关,客户的语气是 frustrated,需要当天处理的概率高达 0.85。

我最看重的是接下来这一步。下一步动作不必靠生成的文字来推断,应用直接拿这些结构化的概率做确定性处理就行:路由投诉,升级处理,置信度低于设定的阈值时就转人工复核。