NingLab/MMECInstruct
Viewer • Updated • 75k • 82
This repo contains the models for "Captions Speak Louder than Images (CASLIE): Generalizing Foundation Models for E-commerce from High-quality Multimodal Instruction Data"
The CASLIE-L model is instruction-tuned from the large base model Llama-2-13b-chat.
@inproceedings{ling2025captions,
title={Captions Speak Louder than Images: Generalizing Foundation Models for E-commerce from High-quality Multimodal Instruction Data},
author={Ling, Xinyi and Du, Hanwen and Peng, Bo and Zhu, Zhihui and Ning, Xia},
booktitle={Proceedings of the 14th International Joint Conference on Natural Language Processing and the 4th Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics},
pages={743--768},
year={2025}
}