decision.host

决策模型

决策模型(Decision Model,也叫 System One 模型或类型化概率决策模型)是一类只做判断、不写文章的模型。给它一段内容(state)和几个类型化的问题答案,它直接返回答案概率。它不生成解释、不做推理链、不输出自由文本。这个品类的起点是 2026 年 9 月 15 日 TypeSafe AI 发布的 Jev 1.13,此后进入了高速发展期。本站会实时更新收录最新的模型、评估、论文等等相关进展。

199模型
304评估集
18,054得分记录
1,197相关论文
85开源权重
30多模态

榜单速览

口径不同,分数不可跨榜换算

S1MB Task Avg(英文文本)

完整 →
1 OpenJev-27B 62.6
2 AutoJev-27B 60.8
3 Eikos 27B 59.9
4 Jev 1.13 59.6

Decision Index Vision(多模态)

完整 →
1 JEV-27B-VL 69.8
5 JPT-9B 61.9
Amir Rafe、Subasish Das · 系统与工程 · stat.AP
Road safety programs count the coded fields of police crash records, while the officer's narrative, which often records factors the fields omit, is rarely read. A safety office thus cannot tell how much its counts miss or where to review. This study develops and evaluates a system that joins both views of the 5,601,890 Texas crashes from 2017 to 2025 into population estimates with stated validity. An in-context tabular foundation model, Kumo Tabular, reads the coded record of every crash, a calibrated System One model, Jev, reads the narratives of two probability samples, and human judgments recalibrate its probabilities. A multiwave predict-then-debias estimator joins the three tiers, and a second human tier drawn with recorded probabilities checks the estimates by design. For hydroplaning, medical episodes, fatigue, animals, and phone use, the narrative documents more injury crashes than the coded field, 15,074 against 7,340 for phone use, and the human check agrees with all fifteen estimates within its margin. A re-read list ranked by Kumo Tabular finds confirmed discordance 7 to 58 times as often as random reading. At the planning cost of human coding, one further round of human judgments would cut the root mean square relative half-width from 22.0 to 16.2 percent, against 21.2 for reading every narrative. Two calibrated readers of different views, joined by a sampling design, give a safety office counts, a discordance map, a validated re-read list, and a reading budget, with Kumo Tabular reading the table at 15 times the speed of TabPFN 3.5.
arXiv 摘要 PDF 26 pages, 9 figures, 8 tables. Code: https://github.com/pozapas/kumo-jev-crash-records Jev / TypeSafe
Baoteng Li、Wenzhuo Wu、Kongming Liang 等 4 人 · 方法与训练 · cs.CV
Multi-subject image generation requires rewards that verify whether requested attributes, actions, and relations hold for the specified reference subjects. Subject presence alone does not establish that the correct subjects participate in a requested interaction. We present reference-bound Visual Jev rewards that turn these visual decisions into generator training signals. Each subject-related question receives a positive label only when the requested condition and the relevant reference identities hold jointly. We construct fixed questions offline, train a Qwen3.5-4B verifier with binary supervision, and directly read Yes probabilities from its language-model head. Their mean supplies a GRPO reward while retaining individual judgments for inspection. Using 200 MICo-150K training tasks and 30 updates, the framework raises a GPT-5.4 composite score from 41.78 to 52.50 on a manually selected 897-task MICo-Bench subset; direct 27B rewards yield 51.84. Each reward is tested in one GRPO run, and offline human evaluation does not establish a statistically significant advantage over direct scoring. The study provides an initial implementation and evaluation of Visual Jev as a reference-bound reward for multi-subject image generation.
arXiv 摘要 PDF Jev / TypeSafe
Wooyoung Jung · 方法与训练 · cs.AI
Artificial intelligence supports building operations in several forms, each with its own barrier. Expert rules must be tuned for every system, supervised models need labeled data that buildings rarely record, and language models return free text that requires human-in-the-loop checking, since their stated confidence is unreliable. A newer kind of pretrained model, here called a decision model, returns a probability for every allowed answer, so one model could serve many decisions without training. This study answers three open questions for fault diagnosis in heating, ventilation, and air-conditioning systems: which decisions such a model can make, what input it needs, and whether its probabilities hold when conditions change. On 128 fault days from four public datasets of real equipment, faults are graded by the reasoning their diagnosis demands, with data given raw, as physical features, or with Brick topology. The decision model Jev, open language models, and a supervised model face nine tests that change season, control configuration, or building. Given physical features, Jev and the larger open model diagnosed faults whose evidence one feature carries, but not faults that need operating context. Under shift they kept their accuracy and calibration, while the supervised model lost 0.33 macro-F1 yet led or tied within a building. Their probabilities still needed correction, and detection was weak. The study maps which faults a decision model can diagnose and from what input, and supports a division of work in which code computes the physics and the model ranks candidate faults for an operator.
arXiv 摘要 PDF 58 pages, 4 figures, 12 tables. Submitted to Energy and Buildings Jev / TypeSafe
Shuyu Gan、Young-Jun Lee、Dongyeop Kang · 方法与训练 · cs.CL
Typed decision models answer a declared question without generating text: a decision head returns a probability for each of the declared options in a single forward pass. A single pass is fast, intuitive System 1 thinking. We study what lies between one pass and generated reasoning: looping, in which the same layers are recursively applied several times before one typed readout. Each loop lets the model revise its hidden state before it commits to an answer, without generating a token; we call this System 1.5 thinking. We propose SanSi, which turns a pre-trained looped language model into a typed decision model. The option probabilities are read after every loop, and every loop is trained with a proper scoring rule, so that one model serves every budget from one loop to eight in a single run. On 10,027 test decisions from 59 sources, SanSi reaches 72.0% accuracy: 13.5 points above a non-looped model of the same shape trained with the same recipe, 5.3 points above a newer non-looped model of its size, and 1.8 points below one with three times the parameters. On two depth-controlled tasks, loops extend the solvable depth beyond the depths seen in training, where the larger single-pass model fails. Used as the judge for policy optimization with reinforcement learning, without gold answers, SanSi raises the generator's F1 by 7.7 points.
arXiv 摘要 PDF 43 pages, 15 figures, 42 tables. Project page: https://minnesotanlp.github.io/Sansi/ 决策模型System 1 双过程类型化决策决策头 / 读出严格评分规则

最新动态

全部 →
2026-10-06
Decision Index 0.3 发布:Perplexity Decider v1.1 登顶,Jev 掉到第 3
新版权重为 0.2×public + 0.5×same_skills + 0.3×new_domains。Perplexity Decider v1.1 (27B) 以 62.8 分第一,Fastino GLiDE 60.2 第二,Jev 1.13 60.1 第三。同版还发布了 Vision 子榜,JEV-27B-VL 以 69.8 分领先。
2026-10-06
Intern-Decision 0.8B/2B/4B 发布:微调语言主干、冻结视觉塔
上海 AI Lab 系的开源多模态决策模型,每字段一个 <decision> token,单张 RTX 4090 约 33 ms/query,训练代码一并开源。
2026-10-01
Cloudflare 开源 Clef / Clef-flash,并推出 RL 微调平台
Clef 27B(Qwen3.8-27B 底座)、Clef-flash 9B(Qwen3.5-9B),Apache-2.0,带视觉编码器、64K 上下文,一轮前向并行给所有选项打分。训练用 label-smoothed CE + Brier loss 做校准,再加 RLCD。这是第一个既有开源权重、又有官方微调服务的决策模型。
2026-10-01
OpenAI 推出 Decisions API(public beta),也是多模态的
端点 POST /v1/decisions,只有 gpt-6-luna 一个模型,支持 text + image 输入,三种问法 predicate / choice / score。输入 $0.10/百万 token,输出免费。支持 ZDR 与 HIPAA,区域限美国与欧洲。

数据来源

每次更新都会记录抓取时间与校验值