Gemma-4-31B-it safety-repair checkpoints: multilingual safety and cybersecurity selectivity. Gated under LSAL v1.2.
AI & ML interests
Frontier research around Safe and aligned intelligence
Recent Activity
View all activity
Papers
Forgetting That Sticks: Quantization-Permanent Unlearning via Circuit Attribution
$C$-$ΔΘ$: Circuit-Restricted Weight Arithmetic for Selective Refusal
Organization Card
Lexsi Labs drives Aligned and Safe AI Frontier Research. Our goal is to build AI systems that are transparent, reliable, and value-aligned, combining interpretability, alignment, and governance to enable trustworthy intelligence at scale.
Research Focus
- Aligned & Safe AI: Frameworks for self-monitoring, interpretable, and alignment-aware systems.
- Explainability & Alignment: Faithful, architecture-agnostic interpretability and value-aligned optimization across tabular, vision, and language models.
- Safe Behaviour Control: Techniques for fine-tuning, pruning, and behavioural steering in large models.
- Risk & Governance: Continuous monitoring, drift detection, and fairness auditing for responsible deployment.
- Tabular & LLM Research: Foundational work on tabular intelligence, in-context learning, and interpretable large language models.
DPO-recovered checkpoints that restore safety to the SafeTune drift checkpoints.
-
Lexsi/gemma3-4b-code-dpo-recover
Image-Text-to-Text • 4B • Updated • 3 -
Lexsi/gemma3-4b-dolly-dpo-recover
Image-Text-to-Text • 4B • Updated • 4 -
Lexsi/gemma3-4b-gsm8k-dpo-recover
Image-Text-to-Text • 4B • Updated • 3 -
Lexsi/llama31-8b-code-dpo-recover
Text Generation • 8B • Updated • 2
Gemma-4-31B-it safety-repair checkpoints: multilingual safety and cybersecurity selectivity. Gated under LSAL v1.2.
DPO-recovered checkpoints that restore safety to the SafeTune drift checkpoints.
-
Lexsi/gemma3-4b-code-dpo-recover
Image-Text-to-Text • 4B • Updated • 3 -
Lexsi/gemma3-4b-dolly-dpo-recover
Image-Text-to-Text • 4B • Updated • 4 -
Lexsi/gemma3-4b-gsm8k-dpo-recover
Image-Text-to-Text • 4B • Updated • 3 -
Lexsi/llama31-8b-code-dpo-recover
Text Generation • 8B • Updated • 2
models 41
Lexsi/Gemma-4-31B-it-Cybersecurity-Safety-Repair
Image-Text-to-Text • 31B • Updated • 156 • 2
Lexsi/Gemma-4-31B-it-Multilingual-Safety-Repair
Image-Text-to-Text • 31B • Updated • 69 • 2
Lexsi/qwen3-4b-gsm8k-dpo-recover
Text Generation • 4B • Updated • 2
Lexsi/qwen3-4b-dolly-dpo-recover
Text Generation • 4B • Updated • 2
Lexsi/qwen3-4b-code-dpo-recover
Text Generation • 4B • Updated • 2
Lexsi/llama32-3b-medical-dpo-recover
Text Generation • 3B • Updated • 2
Lexsi/llama32-3b-legal-dpo-recover
Text Generation • 3B • Updated • 2
Lexsi/llama32-3b-gsm8k-dpo-recover
Text Generation • 3B • Updated • 6
Lexsi/llama32-3b-dolly-dpo-recover
Text Generation • 3B • Updated • 2
Lexsi/llama32-3b-code-dpo-recover
Text Generation • 3B • Updated • 2