Laya fine-tuned for four news questions
Browse files- README.md +109 -0
- model.safetensors +3 -0
- questions.json +48 -0
README.md
ADDED
|
@@ -0,0 +1,109 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
---
|
| 2 |
+
license: apache-2.0
|
| 3 |
+
base_model: convaiinnovations/laya
|
| 4 |
+
language:
|
| 5 |
+
- multilingual
|
| 6 |
+
tags:
|
| 7 |
+
- laya
|
| 8 |
+
- system-one
|
| 9 |
+
- news
|
| 10 |
+
- classification
|
| 11 |
+
- rlcd
|
| 12 |
+
pipeline_tag: text-classification
|
| 13 |
+
---
|
| 14 |
+
|
| 15 |
+
# laya-news-decisions
|
| 16 |
+
|
| 17 |
+
A fine-tuned [Laya](https://huggingface.co/convaiinnovations/laya) (multilingual
|
| 18 |
+
variant, 322M parameters) that answers four typed questions about a news article
|
| 19 |
+
in one forward pass:
|
| 20 |
+
|
| 21 |
+
| key | type | question |
|
| 22 |
+
|---|---|---|
|
| 23 |
+
| `topic` | choice, 10 classes | which category does this news report belong to |
|
| 24 |
+
| `impact` | score, 5 levels | how far do the consequences reach, from one person to global |
|
| 25 |
+
| `violence` | yes/no | does the article describe physical violence |
|
| 26 |
+
| `civilian_harm` | score, 5 levels | how severe are the consequences for civilians |
|
| 27 |
+
|
| 28 |
+
It was trained with [laya-rlcd-training](https://huggingface.co/InfinimindCreations/laya-rlcd-training),
|
| 29 |
+
the open training loop we wrote for Laya. This model is what that loop produces
|
| 30 |
+
on a real corpus.
|
| 31 |
+
|
| 32 |
+
## Use
|
| 33 |
+
|
| 34 |
+
```python
|
| 35 |
+
import json, laya
|
| 36 |
+
from huggingface_hub import hf_hub_download
|
| 37 |
+
from safetensors.torch import load_file
|
| 38 |
+
|
| 39 |
+
repo = "InfinimindCreations/laya-news-decisions"
|
| 40 |
+
agent = laya.load("convaiinnovations/laya", device="cpu", subfolder="multilingual")
|
| 41 |
+
agent.model.load_state_dict(load_file(hf_hub_download(repo, "model.safetensors")), strict=False)
|
| 42 |
+
questions = json.load(open(hf_hub_download(repo, "questions.json")))
|
| 43 |
+
|
| 44 |
+
text = "Floods in Bangladesh displace 40,000 families"
|
| 45 |
+
print(agent.predict({"headline": text}, questions)["answers"])
|
| 46 |
+
```
|
| 47 |
+
|
| 48 |
+
The state key is `headline` because that is the field the model was trained on,
|
| 49 |
+
but we fed it the article body (first 2,800 characters) during training, and
|
| 50 |
+
that is what it expects. Titles alone work, less well.
|
| 51 |
+
|
| 52 |
+
**Use `questions.json` exactly as it is.** The model is sensitive to the wording
|
| 53 |
+
of the options it was trained with, not to the instruction sentence. Two option
|
| 54 |
+
descriptions start with the literal text `NEU v2.`, a leftover editing note from
|
| 55 |
+
our label schema. It is part of the trained wording; removing it changes the
|
| 56 |
+
answers. Renaming the question keys does not (checked on four texts, identical
|
| 57 |
+
output).
|
| 58 |
+
|
| 59 |
+
## Training
|
| 60 |
+
|
| 61 |
+
- About 92,000 news articles from 2026, many languages, labelled on the full text
|
| 62 |
+
by DeepSeek v4.1 Flash.
|
| 63 |
+
- One pass, full fine-tune on one A100, batch 16, G=32, lr 1e-5 with cosine
|
| 64 |
+
decay. Roughly 11 USD.
|
| 65 |
+
- Only four questions on purpose: a seven-question version of the same model was
|
| 66 |
+
worse on four of six shared questions. Every extra question costs the others.
|
| 67 |
+
|
| 68 |
+
## Evaluation
|
| 69 |
+
|
| 70 |
+
176 gold judgements, each made by three independent LLM annotators and kept by
|
| 71 |
+
majority vote. Accuracy is exact match; for the ordinal questions `±1` is the
|
| 72 |
+
share within one level and `ρ` the rank correlation.
|
| 73 |
+
|
| 74 |
+
| question | selected checkpoint | end of run | notes |
|
| 75 |
+
|---|---|---|---|
|
| 76 |
+
| topic | 0.784 | 0.756 | majority class 0.148 |
|
| 77 |
+
| impact | 0.614 | 0.557 | ±1 0.977 · ρ 0.775 |
|
| 78 |
+
| violence | 0.943 | 0.966 | AUC 0.977 |
|
| 79 |
+
| civilian harm | 0.659 | 0.642 | ±1 0.943 · ρ 0.707 |
|
| 80 |
+
|
| 81 |
+
**Read the first column with care.** The released checkpoint was selected as the
|
| 82 |
+
best of 14 evaluations on this same gold set, so its numbers are the optimistic
|
| 83 |
+
end. The second column is the last checkpoint of the run and is closer to what
|
| 84 |
+
you should expect on unseen data. Adjacent evaluations moved by up to 4.5 points on
|
| 85 |
+
topic, which is eight articles out of 176.
|
| 86 |
+
|
| 87 |
+
For scale, the labelling models on the same 176 articles and the same topic
|
| 88 |
+
question: DeepSeek v4.1 Flash 0.835, Gemini 3 Flash 0.812, Claude Haiku 4.5
|
| 89 |
+
0.778. The fine-tuned 322M encoder sits in that range at a fraction of the cost
|
| 90 |
+
and runs locally.
|
| 91 |
+
|
| 92 |
+
Where it is weak: `politics`. It gets half of those right, while its teacher got
|
| 93 |
+
three quarters. Better labels did not fix it (we tried, from scratch and as a
|
| 94 |
+
continuation, both worse), so it is capacity or data volume, not label quality.
|
| 95 |
+
|
| 96 |
+
## Limits
|
| 97 |
+
|
| 98 |
+
- The gold set is small and curated. It over-represents high-impact articles
|
| 99 |
+
compared to a real news stream, so precision in the field will be lower than
|
| 100 |
+
these numbers suggest.
|
| 101 |
+
- Topic categories overlap in the real world (is a sanctions package politics or
|
| 102 |
+
finance?). Even the large labelling models disagree with each other on about
|
| 103 |
+
one article in six.
|
| 104 |
+
- The 0.5 cut-off is not a decision boundary for `violence`. Fit your threshold on
|
| 105 |
+
your own data.
|
| 106 |
+
|
| 107 |
+
## License
|
| 108 |
+
|
| 109 |
+
Apache 2.0, following Laya itself.
|
model.safetensors
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:56b22b357120265cd8a2611638ae25c4ffbb853cf80ba744d9d49df766fd088a
|
| 3 |
+
size 1287653720
|
questions.json
ADDED
|
@@ -0,0 +1,48 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"topic": {
|
| 3 |
+
"type": "choice",
|
| 4 |
+
"instructions": "Which category does this news report belong to?",
|
| 5 |
+
"criteria": {
|
| 6 |
+
"conflict": "armed conflict, war, military operations, attacks, strikes on targets",
|
| 7 |
+
"security": "internal security, police, crime, terrorism response, cybersecurity, intelligence services. NOT everything that merely sounds dangerous.",
|
| 8 |
+
"politics": "government, elections, diplomacy, legislation, parties, court rulings on political matters",
|
| 9 |
+
"finance": "economy, markets, companies, currency, trade, business crime",
|
| 10 |
+
"technology": "technology, IT, AI, science, space",
|
| 11 |
+
"health": "medicine, disease, healthcare system, epidemics",
|
| 12 |
+
"humanitarian": "displacement, hunger, aid operations, human-rights situation, refugee and civilian suffering in a crisis",
|
| 13 |
+
"disaster": "NEU v2. earthquakes, floods, storms with damage, fires, industrial and traffic accidents with casualties. The EVENT itself; the aid response to it is humanitarian.",
|
| 14 |
+
"infrastructure": "NEU v2. power supply, transport networks, water, telecoms, pipelines: outages, failures, maintenance, expansion. Drought as a supply problem belongs here, drought as a famine belongs to humanitarian.",
|
| 15 |
+
"off-topic": "sport, entertainment, celebrities, lifestyle, plain weather forecasts without an event, and everything that is not a real article"
|
| 16 |
+
}
|
| 17 |
+
},
|
| 18 |
+
"impact": {
|
| 19 |
+
"type": "score",
|
| 20 |
+
"instructions": "How far do the consequences of this event reach?",
|
| 21 |
+
"criteria": [
|
| 22 |
+
"single person or single place",
|
| 23 |
+
"one city or one organisation",
|
| 24 |
+
"one country or one sector",
|
| 25 |
+
"several countries or a whole industry",
|
| 26 |
+
"global or systemic"
|
| 27 |
+
]
|
| 28 |
+
},
|
| 29 |
+
"violence": {
|
| 30 |
+
"type": "noul",
|
| 31 |
+
"instructions": "Does this news report describe physical violence against people?",
|
| 32 |
+
"criteria": {
|
| 33 |
+
"true": "killing, wounding, attack, assault, bombardment against people",
|
| 34 |
+
"false": "everything else, including threats, weapons deals and verbal attacks"
|
| 35 |
+
}
|
| 36 |
+
},
|
| 37 |
+
"civilian_harm": {
|
| 38 |
+
"type": "score",
|
| 39 |
+
"instructions": "How severely are civilians affected?",
|
| 40 |
+
"criteria": [
|
| 41 |
+
"civilians not affected",
|
| 42 |
+
"civilians affected indirectly (restrictions, supply)",
|
| 43 |
+
"few civilians injured or displaced",
|
| 44 |
+
"many civilians affected, deaths",
|
| 45 |
+
"mass casualties among civilians"
|
| 46 |
+
]
|
| 47 |
+
}
|
| 48 |
+
}
|