Learn how to attach labels such as sentiment to extracted entity spans. Attributes are span-conditioned: the model first finds entities, then scores attribute labels at those exact spans.
This is different from document-level classification (classify_text). A review can be mixed overall while individual products are positive or negative.
- Setup
- How Attributes Work
- Sentiment on Product Mentions
- Restricting Attributes to Entity Types
- Single-Label vs Multi-Label Attributes
- Multiple Attribute Groups
- Qualifying Labels
- Combining with Other Schema Tasks
- Reading the Output
- Best Practices
GLiNER2.5 is the boundary architecture. Load it with AutoExtractor so the checkpoint's architecture field selects BoundaryExtractor.
from gliner2 import AutoExtractor, AttributeGroup
model = AutoExtractor.from_pretrained("fastino/gliner2.5-multi-v1")Other GLiNER2.5 sizes: gliner2.5-small-v1 (English, fast) and gliner2.5-base-v1 (English). See the README model catalog.
entity_attributes() must be called after entities(). Attribute group names cannot be text, confidence, start, or end.
- You declare content entity types (
product,company, …). - You declare one or more
AttributeGroups (for examplesentiment). - Attribute labels are encoded as internal queries. They are not returned as extra entity types.
- After entity decoding, each retained span is force-scored for every applicable attribute group.
- The chosen label(s) are attached to that entity object.
Single-label groups use softmax (one value). Multi-label groups use sigmoid plus threshold.
text = "The new iPhone camera is excellent, but the battery life is disappointing."
schema = (
model.create_schema()
.entities(["product"])
.entity_attributes({
"sentiment": AttributeGroup(
["positive", "negative", "neutral"],
applies_to=["product"],
qualify_labels=True,
)
})
)
result = model.extract(
text,
schema,
include_spans=True,
include_confidence=True,
)
print(result)
# {
# "entities": {
# "product": [
# {
# "text": "iPhone camera",
# "start": 8,
# "end": 21,
# "confidence": 0.91,
# "sentiment": {"label": "positive", "confidence": 0.87},
# },
# {
# "text": "battery life",
# "start": 35,
# "end": 47,
# "confidence": 0.88,
# "sentiment": {"label": "negative", "confidence": 0.84},
# },
# ]
# }
# }Always request spans while developing so you can check text[start:end] == entity["text"].
for entity in result["entities"]["product"]:
assert text[entity["start"]:entity["end"]] == entity["text"]
print(entity["text"], "→", entity["sentiment"]["label"])applies_to limits which entity types receive the group. Company names in the example below are extracted without a sentiment field.
schema = (
model.create_schema()
.entities(["product", "company"])
.entity_attributes({
"sentiment": AttributeGroup(
["positive", "negative", "neutral"],
applies_to=["product"],
)
})
)
result = model.extract(
"Apple's new AirPods are fantastic.",
schema,
include_spans=True,
include_confidence=True,
)
# company spans have no "sentiment" key
# product spans include sentimentIf applies_to is omitted, the group is scored on every declared entity type. Unknown names in applies_to raise ValueError.
One mutually exclusive value, via softmax. Use this for sentiment, polarity, or status.
AttributeGroup(
["positive", "negative", "neutral"],
multi_label=False, # default
)Without include_confidence, the attribute is still a dict with label (and confidence is omitted from the parent entity, but attribute objects from the span path typically still carry a score when confidence is requested).
Independent sigmoid decisions. Several labels can fire on the same span.
schema = (
model.create_schema()
.entities(["product"])
.entity_attributes({
"aspects": AttributeGroup(
["price", "quality", "design", "battery", "support"],
multi_label=True,
threshold=0.4,
applies_to=["product"],
)
})
)
result = model.extract(
"The Pixel fold is expensive but the build quality and design are outstanding.",
schema,
include_spans=True,
include_confidence=True,
)
# "aspects": [
# {"label": "price", "confidence": 0.71},
# {"label": "quality", "confidence": 0.82},
# {"label": "design", "confidence": 0.79},
# ]Lower threshold to recall more aspects; raise it to reduce noise.
A span can carry several independent groups. Labels must be unique across groups. If a label string would collide with an entity type or another group, set qualify_labels=True.
schema = (
model.create_schema()
.entities({
"product": "Consumer devices or software products",
"company": "Company or brand names",
})
.entity_attributes({
"sentiment": AttributeGroup(
["positive", "negative", "neutral"],
applies_to=["product"],
qualify_labels=True,
),
"urgency": AttributeGroup(
["low", "medium", "high"],
applies_to=["product"],
qualify_labels=True,
),
})
)
result = model.extract(
"The charger failed on day one and we need a replacement immediately.",
schema,
include_spans=True,
include_confidence=True,
)Each product span then includes both sentiment and urgency.
qualify_labels=True prefixes model-facing labels with the group name (sentiment: positive) while returning the short label (positive) in the result. Use it when:
- the same word is both an entity type and an attribute (
status,type,label) - two groups share similar vocabulary
- you want the encoder to see an unambiguous query
AttributeGroup(
["positive", "negative", "neutral"],
qualify_labels=True,
)If an unqualified attribute label collides with an entity type, schema construction raises:
Attribute labels collide with entity labels: ...; use qualify_labels=True
Attributes compose with classification, relations, and structures in one extract call.
schema = (
model.create_schema()
.entities(["product", "company"])
.entity_attributes({
"sentiment": AttributeGroup(
["positive", "negative", "neutral"],
applies_to=["product"],
qualify_labels=True,
)
})
.classification("review_type", ["unboxing", "complaint", "comparison", "praise"])
)
result = model.extract(
"Unboxing the new Surface Laptop: the keyboard is a joy, Microsoft nailed it.",
schema,
include_spans=True,
include_confidence=True,
)Document-level review_type is independent of per-span sentiment.
For documents longer than the model window, use extract_long with the same schema. See Long-Context Extraction.
include_spans |
include_confidence |
Entity object |
|---|---|---|
True |
True |
{text, start, end, confidence, sentiment: {label, confidence}} |
True |
False |
{text, start, end, sentiment: ...} |
False |
True |
{text, confidence, sentiment: ...} |
Single-label attribute value:
entity["sentiment"]["label"] # "positive"
entity["sentiment"]["confidence"] # 0.87Multi-label attribute value:
[item["label"] for item in entity["aspects"]]Attributes are only attached to retained entities. If a mention is dropped by threshold or overlap policy, it is not attributed.
- Call
entities()first, thenentity_attributes(). - Use
applies_toso companies, dates, and locations are not given product sentiment. - Prefer
qualify_labels=Truein multi-task schemas. - Keep attribute vocabularies short and mutually exclusive for single-label groups.
- Request
include_spans=Trueuntil offsets are verified against the source text. - Do not reuse reserved names (
text,start,end,confidence) as group names. - For document-level sentiment of the whole review, use classification instead of (or in addition to) span attributes.
- Boundary checkpoints score attributes at decoded spans; they are not a second NER pass and will not invent extra entity types in the public output.