@george_onx: 最近看到很多将Jev与GLiNER模型进行比较的情况。如果你正在进行基准测试,GLiNER2.5是我们发布的最新版本…
摘要
GLiNER2.5 是一个多语言边界检查点,用于统一的基于模式的信息提取,支持实体、分类、结构化记录、关系和跨度属性在一个模型中。
查看缓存全文
缓存时间: 2026/09/22 07:53
最近看到很多将 Jev 与 GLiNER 模型进行比较的讨论。
如果您正在进行基准测试,GLiNER2.5 是我们发布的最新版本。请用 @fastinoai 标记测试结果,以便我们了解其实际表现。
🤗 模型权重:https://t.co/WFyZXasPqw https://t.co/r18K0bTVUo
fastino/gliner2.5-multi-v1 · Hugging Face
来源:https://huggingface.co/fastino/gliner2.5-multi-v1 Fastino Agent - 通过单个提示微调 GLiNER (https://fastino.ai/lp/gliner)
使用 Fastino Agent 微调并部署 GLiNER2.5 (https://fastino.ai/?utm_source=huggingface) arXiv 论文 (https://arxiv.org/abs/2507.18546) GitHub (https://github.com/fastino-ai/GLiNER2) 关注 @fastinoAI (https://x.com/fastinoAI)
https://huggingface.co/fastino/gliner2.5-multi-v1#gliner25-multi-unified-schema-based-information-extractionGLiNER2.5 Multi:基于统一架构的信息提取
提取实体、分类文本、解析结构化记录、评估跨度属性以及提取关系——所有这些都在一个统一的边界架构中完成。
GLiNER2.5 Multi 是一个多语言边界检查点。它基于 mDeBERTa-v3-base 构建,是您需要在一个模型中跨语言处理实体、分类、记录和关系时的默认选择。使用 AutoExtractor 加载:检查点的 architecture 字段会自动选择 BoundaryExtractor。
通过 Fastino (https://fastino.ai/) 进行微调。加入 Reddit (https://www.reddit.com/r/GLiNER/) 讨论。
- 🎯 一个模型,多种任务:在一个模式中处理实体、分类、结构化记录、关系和跨度属性
- 📐 边界架构:使用稀疏的开始/结束配对而非固定的跨度宽度网格——支持编码窗口内的任意跨度长度
- 🔗 约束解码:使用
Classifier进行跨任务标签约束,使用JointIE处理类型化的实体-关系图 - 💻 本地推理:通过
gliner2[local]支持 CPU、CUDA 或 MPS——无需外部 API
https://huggingface.co/fastino/gliner2.5-multi-v1#gliner25-familyGLiNER2.5 系列
此卡片适用于 fastino/gliner2.5-multi-v1。所有三个检查点共享相同的公共 API。
https://huggingface.co/fastino/gliner2.5-multi-v1#installation安装
pip install "gliner2[local]"
需要 Python 3.10 或更高版本。[local] 额外依赖会引入 PyTorch,以便您可以加载 Hub 检查点。
https://huggingface.co/fastino/gliner2.5-multi-v1#load-the-model加载模型
始终使用 AutoExtractor 加载 GLiNER2.5。GLiNER2.from_pretrained(...) 是旧版 跨度 加载器,不会派发到此检查点。
from gliner2 import AutoExtractor
model = AutoExtractor.from_pretrained("fastino/gliner2.5-multi-v1")
print(type(model).__name__)
print(model.config.architecture)
# BoundaryExtractor
# boundary
可选的设备、fp16 和编译标志:
model = AutoExtractor.from_pretrained(
"fastino/gliner2.5-multi-v1",
map_location="cuda", # 或 "cpu" / "mps"
quantize=True, # GPU 上的 fp16 权重
compile=True, # 首次追踪调用后进行 torch.compile
)
print(type(model).__name__, next(model.parameters()).device)
# BoundaryExtractor cuda:0
https://huggingface.co/fastino/gliner2.5-multi-v1#usage用法
https://huggingface.co/fastino/gliner2.5-multi-v1#entity-extraction实体提取
text = "Apple CEO Tim Cook announced iPhone 15 in Cupertino yesterday."
result = model.extract_entities(
text,
["company", "person", "product", "location"],
include_confidence=True,
include_spans=True,
)
print(result)
# {
# "entities": {
# "company": [{"text": "Apple", "start": 0, "end": 5, "confidence": 0.98}],
# "person": [{"text": "Tim Cook", "start": 10, "end": 18, "confidence": 0.97}],
# "product": [{"text": "iPhone 15", "start": 29, "end": 38, "confidence": 0.96}],
# "location": [{"text": "Cupertino", "start": 42, "end": 51, "confidence": 0.95}],
# }
# }
返回的偏移量是原始字符串中的半开区间字符跨度:text[start:end] == entity["text"]。
当标签具有领域特定性时,添加描述:
result = model.extract_entities(
"Patient received 400mg ibuprofen for severe headache at 2 PM.",
{
"medication": "药物或药品物质的名称",
"dosage": "剂量,如 400mg、2 片或 5ml",
"symptom": "报告的症状或状况",
"time": "时钟时间或相对时间",
},
include_spans=True,
)
print(result)
# {
# "entities": {
# "medication": [{"text": "ibuprofen", "start": 23, "end": 32}],
# "dosage": [{"text": "400mg", "start": 17, "end": 22}],
# "symptom": [{"text": "severe headache", "start": 37, "end": 52}],
# "time": [{"text": "2 PM", "start": 56, "end": 60}],
# }
# }
https://huggingface.co/fastino/gliner2.5-multi-v1#text-classification文本分类
使用 classify_text 进行独立的单任务解码:
result = model.classify_text(
"这台笔记本电脑性能惊人但电池续航太差!",
{"sentiment": ["positive", "negative", "neutral"]},
)
print(result)
# {"sentiment": "negative"}
result = model.classify_text(
"相机画质出色,性能尚可,但电池续航很差。",
{
"aspects": {
"labels": ["camera", "performance", "battery", "display", "price"],
"multi_label": True,
"cls_threshold": 0.4,
}
},
)
print(result)
# {"aspects": ["camera", "performance", "battery"]}
https://huggingface.co/fastino/gliner2.5-multi-v1#constrained-classification约束分类
当一个任务的标签在逻辑上约束另一个任务时,使用 gliner2.classification.Classifier。classify_text 不会强制执行这些规则。
from gliner2.classification import (
Classifier,
ClassificationSchema,
ClassificationConfig,
)
from gliner2.classification import constraints as C
clf = Classifier.from_pretrained("fastino/gliner2.5-multi-v1")
schema = (
ClassificationSchema()
.single("intent", ["read", "write", "delete"])
.multi("effects", ["read_only", "create", "modify", "delete"], min_labels=1)
.constrain(
C.implies(("intent", "delete"), ("effects", "delete")),
C.excludes(("intent", "read"), ("effects", "delete")),
)
)
result = clf.classify("Delete the temporary file from /tmp", schema)
print(result.value("intent"))
print(result.value("effects"))
print(result.feasible)
print(result.to_dict())
# delete
# ['delete']
# True
# {
# "intent": {
# "value": "delete",
# "confidence": 0.93,
# "probabilities": {"read": 0.02, "write": 0.05, "delete": 0.93},
# },
# "effects": {
# "value": ["delete"],
# "confidence": 0.88,
# "probabilities": {
# "read_only": 0.04, "create": 0.03, "modify": 0.05, "delete": 0.88
# },
# },
# "_meta": {"feasible": True, "decoder": "exact"},
# }
预测参数应放在调用的 ClassificationConfig 中,而不是 from_pretrained 中:
result = clf.classify(
"预览报告",
schema,
config=ClassificationConfig(decoder="beam", beam_size=16),
)
print(result.value("intent"), result.value("effects"), result.feasible)
# read ['read_only'] True
https://huggingface.co/fastino/gliner2.5-multi-v1#relation-extraction关系提取
此检查点是使用 enable_relations=True 训练的。独立解码:
text = "Alice works for Acme in Paris."
result = model.extract_relations(
text,
["works_for", "located_in"],
include_spans=True,
include_confidence=True,
)
print(result)
# {
# "relation_extraction": {
# "works_for": [{
# "head": {"text": "Alice", "start": 0, "end": 5, "confidence": 0.91},
# "tail": {"text": "Acme", "start": 16, "end": 20, "confidence": 0.91},
# }],
# "located_in": [{
# "head": {"text": "Acme", "start": 16, "end": 20, "confidence": 0.87},
# "tail": {"text": "Paris", "start": 24, "end": 29, "confidence": 0.87},
# }],
# }
# }
或通过模式:
schema = model.create_schema().relations(
{"works_for": {"threshold": 0.6}, "located_in": {"threshold": 0.6}}
)
result = model.extract(text, schema, include_spans=True)
print(result)
# {
# "relation_extraction": {
# "works_for": [{
# "head": {"text": "Alice", "start": 0, "end": 5},
# "tail": {"text": "Acme", "start": 16, "end": 20},
# }],
# "located_in": [{
# "head": {"text": "Acme", "start": 16, "end": 20},
# "tail": {"text": "Paris", "start": 24, "end": 29},
# }],
# }
# }
独立提取不能保证 works_for 的头部是人员,尾部是组织。
https://huggingface.co/fastino/gliner2.5-multi-v1#joint-information-extraction联合信息提取
JointIE 对提及和关系候选进行评分,然后搜索具有类型化端点和唯一性约束的全局一致图。
from gliner2.joint_ie import JointIE, JointIEConfig
joint = JointIE.from_pretrained("fastino/gliner2.5-multi-v1")
schema = (
joint.create_schema()
.entities(["person", "organization", "location"])
.relation("works_for", "person", "organization", unique_head=True)
.relation("located_in", "organization", "location")
.no_self_loops()
)
result = joint.extract(
"Alice works for Acme in Paris. Bob joined Acme last year.",
schema,
config=JointIEConfig(optimizer="beam", beam_size=32),
)
print(result.feasible)
print(result.to_dict())
# True
# {
# "entities": [
# {"id": "e1", "type": "person", "text": "Alice", "start": 0, "end": 5, "confidence": 0.94},
# {"id": "e2", "type": "organization", "text": "Acme", "start": 16, "end": 20, "confidence": 0.92},
# {"id": "e3", "type": "location", "text": "Paris", "start": 24, "end": 29, "confidence": 0.90},
# {"id": "e4", "type": "person", "text": "Bob", "start": 31, "end": 34, "confidence": 0.91},
# ],
# "relations": [
# {"type": "works_for", "head": "e1", "tail": "e2", "confidence": 0.88},
# {"type": "works_for", "head": "e4", "tail": "e2", "confidence": 0.81},
# {"type": "located_in", "head": "e2", "tail": "e3", "confidence": 0.86},
# ],
# }
始终检查 result.feasible。False 表示硬约束无法满足(这与“文本不包含事实”不同)。
for rel in result.relations:
head = result.entity(rel.head)
tail = result.entity(rel.tail)
print(f"{head.text} -{rel.type}-> {tail.text}")
# Alice -works_for-> Acme
# Bob -works_for-> Acme
# Acme -located_in-> Paris
https://huggingface.co/fastino/gliner2.5-multi-v1#span-attributes-people-with-sentiment跨度属性:带有情感的人
属性是基于跨度的。模型首先找到实体,然后在这些精确跨度上评估属性标签。它们不是额外的实体类型,也不是文档级分类。
from gliner2 import AutoExtractor, AttributeGroup
model = AutoExtractor.from_pretrained("fastino/gliner2.5-multi-v1")
text = (
"Alice was delighted with the promotion, "
"but Bob sounded frustrated about the delay."
)
schema = (
model.create_schema()
.entities(["person"])
.entity_attributes({
"sentiment": AttributeGroup(
["positive", "negative", "neutral"],
applies_to=["person"],
qualify_labels=True,
)
})
)
result = model.extract(
text,
schema,
include_spans=True,
include_confidence=True,
)
print(result)
# {
# "entities": {
# "person": [
# {
# "text": "Alice",
# "start": 0,
# "end": 5,
# "confidence": 0.96,
# "sentiment": {"label": "positive", "confidence": 0.89},
# },
# {
# "text": "Bob",
# "start": 44,
# "end": 47,
# "confidence": 0.95,
# "sentiment": {"label": "negative", "confidence": 0.84},
# },
# ]
# }
# }
applies_to=["person"] 可防止情感应用于其他实体类型。qualify_labels=True 将模型面向的查询编码为 sentiment: positive,同时返回简短标签 positive。
限制情感只应用于人员,同时仍提取公司:
schema = (
model.create_schema()
.entities(["person", "organization"])
.entity_attributes({
"sentiment": AttributeGroup(
["positive", "negative", "neutral"],
applies_to=["person"],
qualify_labels=True,
)
})
)
result = model.extract(
"Alice praised Microsoft, but Bob criticized OpenAI.",
schema,
include_spans=True,
include_confidence=True,
)
print(result)
# {
# "entities": {
# "person": [
# {
# "text": "Alice",
# "start": 0,
# "end": 5,
# "confidence": 0.96,
# "sentiment": {"label": "positive", "confidence": 0.88},
# },
# {
# "text": "Bob",
# "start": 29,
# "end": 32,
# "confidence": 0.95,
# "sentiment": {"label": "negative", "confidence": 0.86},
# },
# ],
# "organization": [
# {"text": "Microsoft", "start": 14, "end": 23, "confidence": 0.97},
# {"text": "OpenAI", "start": 44, "end": 50, "confidence": 0.96},
# ],
# }
# }
组织跨度没有 sentiment 字段。人员跨度有。
https://huggingface.co/fastino/gliner2.5-multi-v1#structured-records结构化记录
记录模式保留实例身份(谁购买了什么),而不是将字段展平为不相关的列表。使用锚点字段启用 natural 模式:
schema = (
model.create_schema()
.structure("purchase", mode="natural", anchor="buyer")
.field("buyer", dtype="str", cardinality="required_one")
.field("item", dtype="str", cardinality="required_one")
)
result = model.extract(
"Alice bought apples and Bob bought oranges.",
schema,
)
print(result)
# {
# "purchase": [
# {"buyer": "Alice", "item": "apples"},
# {"buyer": "Bob", "item": "oranges"},
# ]
# }
此检查点是使用 enable_records=True 训练的。
https://huggingface.co/fastino/gliner2.5-multi-v1#task-combination任务组合
在一次 extract 调用中组合实体、跨度属性、分类、关系和结构:
from gliner2 import AttributeGroup
schema = (
model.create_schema()
.entities({
"person": "Named people",
"organization": "Companies or teams",
"product": "Named products or services",
})
.entity_attributes({
"sentiment": AttributeGroup(
["positive", "negative", "neutral"],
applies_to=["person"],
qualify_labels=True,
)
})
.classification("topic", ["technology", "business", "sports", "politics"])
.relations(["works_for", "announced"])
.structure("announcement", mode="natural", anchor="product")
.field("company", dtype="str")
.field("product", dtype="str", cardinality="required_one")
)
text = "Apple CEO Tim Cook unveiled the iPhone 15 Pro for $999."
result = model.extract(text, schema, include_spans=True, include_confidence=True)
print(result)
# {
# "entities": {
# "person": [{
# "text": "Tim Cook",
# "start": 10,
# "end": 18,
# "confidence": 0.97,
# "sentiment": {"label": "positive", "confidence": 0.82},
# }],
# "organization": [{"text": "Apple", "start": 0, "end": 5, "confidence": 0.98}],
# "product": [{"text": "iPhone 15 Pro", "start":
相似文章
GLiNER2: 一个高效的多任务信息提取系统 具有模式驱动接口
GLiNER2 是一个基于Transformer的统一框架,用于多种NLP任务,通过支持命名实体识别、文本分类和层次数据提取,在单个模型中提供比大型语言模型更高的效率和可访问性。
@mervenoyann: GLM-5.2 与 Opus 4.8 相当,具有 1M 上下文 > 新的 IS 注意力每 4 个稀疏层重用一次索引器(2.9× 每…)
GLM-5.2 是一款可与 Opus 4.8 相媲美的新模型,具有 1M 上下文、新的 IS 注意力机制、改进的推测解码和灵活的思考努力级别。它已在 MIT 许可证下发布,并在 transformers、vLLM 和 SGLang 中提供 Day-0 支持。
GLiNER-Relex:联合命名实体识别与关系提取的统一框架
GLiNER-Relex 是一个用于联合命名实体识别(NER)与关系提取(RE)的统一框架,利用共享的 Transformer 编码器实现零样本能力。该论文展示了模型在标准基准测试中具有竞争力的性能,并将其作为开源 Python 包发布。
对 Google Embeddings 2 与开源模型在多语言稠密检索和 RAG 系统中的基准测试
本文对 Google Embeddings 2 与五个开源模型在多语言稠密检索和 RAG 系统中进行了基准测试,发现 GE2 在准确性上表现最佳但速度较慢,而 mE5-L 作为低延迟的竞争性替代方案。
@MiaAI_lab: GLM-5.2 是目前最好的中文开源模型。输出质量极高——我能真切感受到差异。问题是……
GLM-5.2 被称赞为目前输出质量最好的中文开源模型,但要注意其较高的 token 消耗。用户希望能在 3 台 DGX Sparks 上运行它。