Cytrex commited on
Commit
6543301
·
verified ·
1 Parent(s): ae20735

Laya fine-tuned for four news questions

Browse files
Files changed (3) hide show
  1. README.md +109 -0
  2. model.safetensors +3 -0
  3. questions.json +48 -0
README.md ADDED
@@ -0,0 +1,109 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ license: apache-2.0
3
+ base_model: convaiinnovations/laya
4
+ language:
5
+ - multilingual
6
+ tags:
7
+ - laya
8
+ - system-one
9
+ - news
10
+ - classification
11
+ - rlcd
12
+ pipeline_tag: text-classification
13
+ ---
14
+
15
+ # laya-news-decisions
16
+
17
+ A fine-tuned [Laya](https://huggingface.co/convaiinnovations/laya) (multilingual
18
+ variant, 322M parameters) that answers four typed questions about a news article
19
+ in one forward pass:
20
+
21
+ | key | type | question |
22
+ |---|---|---|
23
+ | `topic` | choice, 10 classes | which category does this news report belong to |
24
+ | `impact` | score, 5 levels | how far do the consequences reach, from one person to global |
25
+ | `violence` | yes/no | does the article describe physical violence |
26
+ | `civilian_harm` | score, 5 levels | how severe are the consequences for civilians |
27
+
28
+ It was trained with [laya-rlcd-training](https://huggingface.co/InfinimindCreations/laya-rlcd-training),
29
+ the open training loop we wrote for Laya. This model is what that loop produces
30
+ on a real corpus.
31
+
32
+ ## Use
33
+
34
+ ```python
35
+ import json, laya
36
+ from huggingface_hub import hf_hub_download
37
+ from safetensors.torch import load_file
38
+
39
+ repo = "InfinimindCreations/laya-news-decisions"
40
+ agent = laya.load("convaiinnovations/laya", device="cpu", subfolder="multilingual")
41
+ agent.model.load_state_dict(load_file(hf_hub_download(repo, "model.safetensors")), strict=False)
42
+ questions = json.load(open(hf_hub_download(repo, "questions.json")))
43
+
44
+ text = "Floods in Bangladesh displace 40,000 families"
45
+ print(agent.predict({"headline": text}, questions)["answers"])
46
+ ```
47
+
48
+ The state key is `headline` because that is the field the model was trained on,
49
+ but we fed it the article body (first 2,800 characters) during training, and
50
+ that is what it expects. Titles alone work, less well.
51
+
52
+ **Use `questions.json` exactly as it is.** The model is sensitive to the wording
53
+ of the options it was trained with, not to the instruction sentence. Two option
54
+ descriptions start with the literal text `NEU v2.`, a leftover editing note from
55
+ our label schema. It is part of the trained wording; removing it changes the
56
+ answers. Renaming the question keys does not (checked on four texts, identical
57
+ output).
58
+
59
+ ## Training
60
+
61
+ - About 92,000 news articles from 2026, many languages, labelled on the full text
62
+ by DeepSeek v4.1 Flash.
63
+ - One pass, full fine-tune on one A100, batch 16, G=32, lr 1e-5 with cosine
64
+ decay. Roughly 11 USD.
65
+ - Only four questions on purpose: a seven-question version of the same model was
66
+ worse on four of six shared questions. Every extra question costs the others.
67
+
68
+ ## Evaluation
69
+
70
+ 176 gold judgements, each made by three independent LLM annotators and kept by
71
+ majority vote. Accuracy is exact match; for the ordinal questions `±1` is the
72
+ share within one level and `ρ` the rank correlation.
73
+
74
+ | question | selected checkpoint | end of run | notes |
75
+ |---|---|---|---|
76
+ | topic | 0.784 | 0.756 | majority class 0.148 |
77
+ | impact | 0.614 | 0.557 | ±1 0.977 · ρ 0.775 |
78
+ | violence | 0.943 | 0.966 | AUC 0.977 |
79
+ | civilian harm | 0.659 | 0.642 | ±1 0.943 · ρ 0.707 |
80
+
81
+ **Read the first column with care.** The released checkpoint was selected as the
82
+ best of 14 evaluations on this same gold set, so its numbers are the optimistic
83
+ end. The second column is the last checkpoint of the run and is closer to what
84
+ you should expect on unseen data. Adjacent evaluations moved by up to 4.5 points on
85
+ topic, which is eight articles out of 176.
86
+
87
+ For scale, the labelling models on the same 176 articles and the same topic
88
+ question: DeepSeek v4.1 Flash 0.835, Gemini 3 Flash 0.812, Claude Haiku 4.5
89
+ 0.778. The fine-tuned 322M encoder sits in that range at a fraction of the cost
90
+ and runs locally.
91
+
92
+ Where it is weak: `politics`. It gets half of those right, while its teacher got
93
+ three quarters. Better labels did not fix it (we tried, from scratch and as a
94
+ continuation, both worse), so it is capacity or data volume, not label quality.
95
+
96
+ ## Limits
97
+
98
+ - The gold set is small and curated. It over-represents high-impact articles
99
+ compared to a real news stream, so precision in the field will be lower than
100
+ these numbers suggest.
101
+ - Topic categories overlap in the real world (is a sanctions package politics or
102
+ finance?). Even the large labelling models disagree with each other on about
103
+ one article in six.
104
+ - The 0.5 cut-off is not a decision boundary for `violence`. Fit your threshold on
105
+ your own data.
106
+
107
+ ## License
108
+
109
+ Apache 2.0, following Laya itself.
model.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:56b22b357120265cd8a2611638ae25c4ffbb853cf80ba744d9d49df766fd088a
3
+ size 1287653720
questions.json ADDED
@@ -0,0 +1,48 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "topic": {
3
+ "type": "choice",
4
+ "instructions": "Which category does this news report belong to?",
5
+ "criteria": {
6
+ "conflict": "armed conflict, war, military operations, attacks, strikes on targets",
7
+ "security": "internal security, police, crime, terrorism response, cybersecurity, intelligence services. NOT everything that merely sounds dangerous.",
8
+ "politics": "government, elections, diplomacy, legislation, parties, court rulings on political matters",
9
+ "finance": "economy, markets, companies, currency, trade, business crime",
10
+ "technology": "technology, IT, AI, science, space",
11
+ "health": "medicine, disease, healthcare system, epidemics",
12
+ "humanitarian": "displacement, hunger, aid operations, human-rights situation, refugee and civilian suffering in a crisis",
13
+ "disaster": "NEU v2. earthquakes, floods, storms with damage, fires, industrial and traffic accidents with casualties. The EVENT itself; the aid response to it is humanitarian.",
14
+ "infrastructure": "NEU v2. power supply, transport networks, water, telecoms, pipelines: outages, failures, maintenance, expansion. Drought as a supply problem belongs here, drought as a famine belongs to humanitarian.",
15
+ "off-topic": "sport, entertainment, celebrities, lifestyle, plain weather forecasts without an event, and everything that is not a real article"
16
+ }
17
+ },
18
+ "impact": {
19
+ "type": "score",
20
+ "instructions": "How far do the consequences of this event reach?",
21
+ "criteria": [
22
+ "single person or single place",
23
+ "one city or one organisation",
24
+ "one country or one sector",
25
+ "several countries or a whole industry",
26
+ "global or systemic"
27
+ ]
28
+ },
29
+ "violence": {
30
+ "type": "noul",
31
+ "instructions": "Does this news report describe physical violence against people?",
32
+ "criteria": {
33
+ "true": "killing, wounding, attack, assault, bombardment against people",
34
+ "false": "everything else, including threats, weapons deals and verbal attacks"
35
+ }
36
+ },
37
+ "civilian_harm": {
38
+ "type": "score",
39
+ "instructions": "How severely are civilians affected?",
40
+ "criteria": [
41
+ "civilians not affected",
42
+ "civilians affected indirectly (restrictions, supply)",
43
+ "few civilians injured or displaced",
44
+ "many civilians affected, deaths",
45
+ "mass casualties among civilians"
46
+ ]
47
+ }
48
+ }