Instructions to use facebook/meta-encoder with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use facebook/meta-encoder with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("feature-extraction", model="facebook/meta-encoder")# pip install -U transformers accelerate # Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("facebook/meta-encoder") model = AutoModelForMultimodalLM.from_pretrained("facebook/meta-encoder", device_map="auto") - Notebooks
- Google Colab
- Kaggle
MetaEncoder-30B
Multimodal System One Encoder with a natural language interface for instruction-following encoding: give it a task and a list of candidates, and it scores the candidates by task specification. Both a task and each candidate are expressed in natural language — instructions, queries, questions, state descriptions, criteria — with image and video as side information.
Key features:
- Natural language interface. No rigid schemas. Describe the task and the candidates in free-form text.
- Highly-efficient cacheable representations. Candidates and task are encoded separately by prompt instructions that ground each other.
- Scales to arbitrarily large candidate sets. Matching is an inner product over representations, so an ANN index can serve millions of candidates without re-running the model.
The model is built by contrastively fine-tuning
meta-models/Muse-Glimmer-30B.
Results
| Benchmark | Score | Metric | Modality | Task Type |
|---|---|---|---|---|
| JEVBench (orig/easy/hard) | 0.9444/1.0000/0.7387 | Accuracy | Text | Closed set |
| ImaJEV (dev/cal) | 0.8555/0.9012 | Accuracy | Multimodal | Closed set |
| MMLU* | 0.7534 | Hit@1 | Text | Closed set |
| MMMU* | 0.5774 | Hit@1 | Multimodal | Closed set |
| Video MMMU (64 frames) | 0.5900 | Hit@1 | Multimodal | Closed set |
| MVBench | 0.5537 | Hit@1 | Multimodal | Closed set |
| NaturalBench | 0.8130 | Hit@1 | Multimodal | Closed set |
| TempCompass | 0.7346 | Hit@1 | Multimodal | Closed set |
| NanoBEIR | 0.6634 | NDCG_linear@10 | Text | Open set |
| MMEB-V3 (Image/Video/VisDoc) | 0.7897/0.6046/0.8138 | Hit@1 / Hit@1 / NDCG@5 | Multimodal | Open set |
Closed set refers to candidate options < 256.
* The option list is added into the task prompt.
Quick start
pip install "transformers>=5.15" torch "torchvision<0.27" accelerate pillow
modeling_metaencoder.py lives in this repo rather than in a package, so fetch the repo once
and put it on your path:
import sys
from huggingface_hub import snapshot_download
path = snapshot_download("facebook/meta-encoder")
sys.path.insert(0, path)
Example
from PIL import Image
from modeling_metaencoder import MetaEncoder
model = MetaEncoder.from_pretrained(path)
# Both a task and a candidate can be multimodal
results = model.match(
task={
"image": Image.open("reference_jacket.jpg"),
"text": "Which of the following media has the same style jacket but in red, with a hood",
},
candidates=[
{"image": Image.open("p1.jpg"), "text": "navy windbreaker, no hood"},
{"image": Image.open("p2.jpg"), "text": "red hooded parka"},
{"image": Image.open("p3.jpg"), "text": "black leather biker jacket"},
{"video": runway_clip, "text": "autumn outerwear runway segment"},
],
)
best = results[0]
best.index, round(best.score, 4), best.candidate
match returns Match(index, score, candidate), best first. Both sides accept text, images
and video in any combination:
- The task's
textcarries the question, query, state, instruction — it is what grounds the task to the candidates. - Each candidate's
textgrounds it to the task, and instruction can also be added into the text to control how that candidate is encoded.
Another example
action = model.match(
task={
"image": Image.open("screen.png"),
"text": "Goal: cancel the pending order. Choose the next UI action.",
},
candidates=[
{"text": "click the 'Orders' tab in the left sidebar"},
{"text": "click the red 'Cancel order' button"},
{"text": "scroll down to the shipping details"},
{"text": "close the dialog without saving"},
],
top_k=1,
)[0]
action.candidate["text"], round(action.score, 4)
Prompt format used in evaluation
For prompt with closed set, it is always recommended to add the options available in prompt, which improves MMLU and MMMU:
{state}
Task: {instruction}
Criteria:
- {label}: {what the label means}
- {label}: {what the label means}
For candidate, there are two scenarios. For the JEV benchmarks, the candidates are the bare labels, which act as pointers into the criteria block. For other datasets, we keep the description in the candidates:
candidates: "white, glycolytic, slow contracting."
"red, oxidative, fast contracting."
"red, oxidative, slow contracting."
...
Reusing a candidate pool
match encodes candidates on every call. For a fixed corpus, embed it once and pass the
tensor back — the scores are identical, and only the task is encoded per query:
index = model.encode(corpus) # once
hits = model.match({"text": question}, index, top_k=5)
corpus[hits[0].index]
→ USAGE.md covers the full task/candidate spec, instructions, video vs. multi-image, the exact prompt format, and troubleshooting.
Serving with vLLM
Text and image inputs can be served with vLLM on its
Transformers backend. Video is not supported there; encode video with MetaEncoder as above.
pip install "vllm==0.23.0" "transformers==5.15.1" "torchvision<0.27"
import os, sys
os.environ["VLLM_ENABLE_V1_MULTIPROCESSING"] = "0" # the model class below is registered in this process
import numpy as np
from huggingface_hub import snapshot_download
from PIL import Image
from transformers import AutoProcessor
from vllm import LLM
from vllm.config import PoolerConfig
path = snapshot_download("facebook/meta-encoder")
sys.path.insert(0, path)
import vllm_metaencoder
vllm_metaencoder.register()
processor = AutoProcessor.from_pretrained(path)
llm = LLM(
model=path, runner="pooling", dtype="bfloat16", max_model_len=8192,
enforce_eager=True, # torch.compile does not handle this model's patched modules yet
limit_mm_per_prompt={"image": 2}, attention_backend="FLASH_ATTN",
pooler_config=PoolerConfig(pooling_type="LAST", use_activation=True), # last token, L2-normalised
)
task = {"text": "Which photo shows a bicycle?"}
candidates = [{"image": Image.open("a.jpg")}, {"text": "a red bicycle leaning on a wall"}]
outputs = llm.embed(
[vllm_metaencoder.to_prompt(processor, item) for item in [task, *candidates]],
tokenization_kwargs=vllm_metaencoder.TOKENIZATION_KWARGS,
)
emb = np.stack([o.outputs.embedding for o in outputs])
scores = emb[1:] @ emb[0] # cosine similarity of each candidate to the task
vllm_metaencoder.py ships in this repo and is required: vLLM's generic backend drops the
weight-less norm inside MuseGlimmer's token embedding and resizes its other weight-less norms,
which register() restores, and to_prompt() builds the exact MetaEncoder prompt with a single
BOS token.
Parity against MetaEncoder (bf16, one GPU): cosine similarity 0.9998 on text and 0.993 on images;
on 200 ImageNet-1K images ranked against all 1,000 labels, 85.5% top-1 with vLLM vs 84.5% with
MetaEncoder, with the same prediction on 97.5% of images. Tested offline on a single GPU with
the settings above; multi-GPU and vllm serve have not been verified yet.
License and use
Licensed under Apache 2.0, inherited from the base model. Use is additionally subject to the base model's Usage Policy. See NOTICE for a statement of what was changed relative to the base model.
Citation
@misc{metaencoder_30b,
title = {MetaEncoder-30B},
note = {Fine-tune of meta-models/Muse-Glimmer-30B},
year = {2026}
}
- Downloads last month
- 6
Model tree for facebook/meta-encoder
Base model
meta-models/Muse-Glimmer-30B