Instructions to use OpenMed/LFM2-1.2B-Longevity-ONNX with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers.js
How to use OpenMed/LFM2-1.2B-Longevity-ONNX with Transformers.js:
// npm i @huggingface/transformers import { pipeline } from '@huggingface/transformers'; // Allocate pipeline const pipe = await pipeline('text-generation', 'OpenMed/LFM2-1.2B-Longevity-ONNX');
LFM2-1.2B-Longevity - ONNX 4-bit
An ONNX conversion of LiquidAI/LFM2-1.2B-Longevity for ONNX Runtime, quantized to a 4-bit body with 8-bit sensitive layers, for on-device use on Android, desktop and the browser (Transformers.js). The model is Liquid AI and Insilico Medicine's Longevity-LLM fine-tune of LFM2; OpenMed made and published this conversion and is not affiliated with or endorsed by the model's authors.
Size and fidelity
| Source (BF16) | This repo | |
|---|---|---|
| Parameters | 1.17 B | 1.17 B (unchanged) |
| Weights, measured from the tensors | - | 0.83 GiB |
| Bits per weight, measured from the tensors | 16 | 6.08 |
Measured against the source model in FP32 on 4,092 tokens of public-domain prose (Project Gutenberg), ONNX Runtime 1.30.0, CPU:
| Metric | Value |
|---|---|
| Mean KL divergence of next-token distributions | 0.053 |
| Top-1 next-token agreement | 85.8% |
| Perplexity change | +2.2% |
The FP32 export this build was quantized from matches the source model (top-1 agreement 100%), so the difference above is the quantization. The width is higher than the "4-bit" label suggests because the tensors that lose most at 4 bits, and the shared embedding table, are kept at 8 bits. Weights exclude the rotary-position tables stored in the graph.
Quantization
| Field | Value |
|---|---|
| Body | int4 asymmetric round-to-nearest (uint4 + packed zero point), block 32, fp32 scales; MatMulNBits accuracy_level 4 |
| Kept at 8 bits | int8 asymmetric, block 32, on the tensors llama.cpp's Q4_K_M rule upgrades (down_proj in the first and last eighth of layers and every third layer between; v_proj by the same rule over attention layers) |
| Embedding and output head | one int8 asymmetric block-32 table stored once, read by GatherBlockQuantized for the embedding and by MatMulNBits (bits 8) for the output head |
| Calibration | none (round-to-nearest); no calibration data |
| Export | Liquid4All/onnx-export (Liquid’s official exporter), FP32 graph |
| Quantizer | onnxruntime 1.30.0 MatMulNBitsQuantizer |
The tokenizer, chat template and generation defaults are the upstream files, unchanged. Exact commands, tool versions and every file's SHA-256 are in openmed_build.json; the graph's metadata_props record the source revision and the modification notice.
Architecture
| Field | Value |
|---|---|
| Source model type | lfm2 (Lfm2ForCausalLM) |
| Design | Hybrid Liquid model: gated short convolutions with 6 grouped-query attention layers out of 16 |
| Hidden size | 2048 |
| Layers | 16 (6 attention, 10 convolution) |
| Vocabulary | 65,536, tied input/output embeddings |
Quick start
Transformers.js (browser or Node)
import { pipeline } from "@huggingface/transformers";
const generator = await pipeline("text-generation", "OpenMed/LFM2-1.2B-Longevity-ONNX", { dtype: "q4" });
const messages = [
{ role: "user", content: "Which biomarkers in a routine blood panel say most about biological age, and why? /no_think" },
];
const output = await generator(messages, { max_new_tokens: 256, do_sample: false });
console.log(output[0].generated_text.at(-1).content);
ONNX Runtime (Python, Android and elsewhere)
The graph is a standard ONNX Runtime decoder: feed input_ids, attention_mask, position_ids and the past_conv.* / past_key_values.* cache inputs, then feed each present* output back as the matching past* input on the next step. It uses ONNX Runtime's com.microsoft operators (MatMulNBits, GatherBlockQuantized, GroupQueryAttention), so run it with ONNX Runtime 1.30 or later (onnxruntime-android on Android).
Tested with onnxruntime 1.30.0 (CPU execution provider) on macOS and with Transformers.js 4.3.0, which produced identical greedy tokens. It has not yet been measured on an Android device.
File set
| File | Size | SHA-256 |
|---|---|---|
LICENSE |
0.0 MiB | 4d28ca14dedc0b3d… |
chat_template.jinja |
0.0 MiB | 013eed60546434b6… |
config.json |
0.0 MiB | 5471c83e87909ee1… |
generation_config.json |
0.0 MiB | 85fa3172f3838eef… |
onnx/model_q4.onnx |
0.2 MiB | e234b7305d77075c… |
onnx/model_q4.onnx_data |
879.7 MiB | 3181f300cc1f445b… |
tokenizer.json |
4.5 MiB | df1d8d5ec5d091b4… |
tokenizer_config.json |
0.0 MiB | a221c5e25a01fb4a… |
Intended use
For research and education on aging biology. It is not a medical device and not a substitute for professional medical advice, diagnosis or treatment. Outputs can be wrong; verify them.
Licence
Distributed under the LFM Open License v1.0, the source model's licence. The ONNX graphs and weight files in onnx/ are modified files: OpenMed converted and quantized them from the source model's PyTorch weights on 2026-09-23. All other files are unchanged upstream copies.
Source: LiquidAI/LFM2-1.2B-Longevity at revision 9b4926c163b8996311827b7534295c8fe84703cf. Please cite the original model when you use this conversion.
- Downloads last month
- 338