← 全部作品All work

MiniCPM5-2B-Jev

一个 2B 参数的「系统一」决策模型:给它一段上下文和一组结构化问题,一次前向计算就返回每个问题经过校准的概率分布。A 2B-parameter System 1 decision model: give it a context and a set of typed questions, and it returns a calibrated probability distribution for every question in a single forward pass.

基于 openbmb/MiniCPM5-2B · LoRA 微调 · 开放权重(Apache-2.0)· 支持 Ollama 和 /v1/systemone 接口Built on openbmb/MiniCPM5-2B · LoRA fine-tune · Open weights (Apache-2.0) · Ollama & /v1/systemone

78.8%
JevBench 总分,≤2B 开放权重模型中第一*JevBench overall — #1 among open-weight models ≤2B*
86.4%
decision-v7 准确率,校准误差 ECE 仅 0.031decision-v7 accuracy, with an ECE of just 0.031
2.0B
参数量,一次前向给出所有问题的结果parameters, answering every question in one forward pass

特点Highlights

一次前向,不生成文字One pass, no generation

直接读出每个选项的概率,不做逐字生成,速度快,结果稳定。Reads out option probabilities directly instead of generating text — fast and deterministic.

概率经过校准Calibrated probabilities

类别平衡的先验去偏加上按题型的温度缩放,输出的概率可以直接当作置信度使用。Class-balanced prior debiasing and per-type temperature scaling make the probabilities usable as confidence.

三种题型Three question types

支持 choice、noul 和 score 三种结构化题型,一个请求里可以同时问多个问题。Supports choice, noul and score questions — ask many at once in a single request.

问题之间互不影响Isolated questions

多个问题共享同一段上下文时只编码一次;增加或调换问题,不会改变其他问题的结果。A shared context is encoded once, and adding or reordering questions never changes another question's answer.

开箱即用Ready to run

一条命令用 Ollama 拉取运行;也可以用自带的 HTTP 服务,在 CUDA、Apple 芯片(MPS)或 CPU 上提供 /v1/systemone 接口。Pull and run it with Ollama in one command, or use the built-in server to expose /v1/systemone on CUDA, Apple silicon (MPS) or CPU.

完整可复现Fully reproducible

训练脚本、数据构建和 7 套基准测评全部开源,测评数据随仓库提供。Training, data building and all seven benchmark suites are open source, with the evaluation data included.

JevBench 对比JevBench comparison

模型Model参数量Params开放权重Open weights总分Overall
MiniCPM5-2B-Jev2.0B✓78.8%
decider-2b v111.9B✓76.2%
Kev-4B (r10)4.2B✓75.8%
system-one-open~2B✓73.2%
system-one (Qwen3-8B)8.2B✓71.9%
Bespoke Nimble 9B9.0B—67.5%

* 在官方 JevBench(231 道决策题)上测评,对比模型取自 2026 年 9 月的公开排行榜。完整的 7 套基准结果见 GitHub。* Evaluated on the official JevBench (231 decisions); other models are from the public leaderboard as of September 2026. Full results across all seven suites are on GitHub.

快速开始Quickstart

用 Ollama 运行(最简单)With Ollama (easiest)

ollama pull actbro/minicpm5-2b-jev:2b

curl http://localhost:11434/v1/systemone -d '{
  "model": "actbro/minicpm5-2b-jev:2b",
  "state": "Ticket #402: Customer wants to cancel and get a full refund.",
  "questions": {
    "is_refund": { "type": "noul", "instructions": "Is the customer asking for a refund?" }
  }
}'

需要 Ollama 0.35 或更高版本,模型约 5 GB。更多用法(Python SDK、/api/generate 兼容模式)见 Ollama 页面。Requires Ollama 0.35 or later; the model is about 5 GB. See the Ollama page for the Python SDK and /api/generate compatibility mode.

从源码运行From source

git clone https://github.com/yuting-ai/minicpm5-2b-jev
cd minicpm5-2b-jev
pip install -r requirements.txt
python serve.py --checkpoint checkpoints/stage2 --port 8013

启动后向 http://127.0.0.1:8013/v1/systemone 发送 JSON 请求即可。Python 调用、测评和训练方法见 README。Then send JSON requests to http://127.0.0.1:8013/v1/systemone. See the README for the Python API, evaluation and training.

开始使用Get started

在 Hugging Face 获取权重Get the weights on Hugging Face

在 Ollama 运行Run with Ollama · 源代码和完整测评结果Source code and full benchmark results