arXiv · 2305.07019
Musketeer: Joint Training for Multi-task Vision Language Model with Task Explanation Prompts
Abstract
We present a vision-language model whose parameters are jointly trained on all tasks and fully shared among multiple heterogeneous tasks which may interfere with each other, resulting in a single model which we named Musketeer. The integration of knowledge across heterogeneous tasks is enabled by a novel feature called Task Explanation Prompt (TEP). With rich and structured information such as task input/output format, TEP reduces interference among tasks, allowing the model to focus on their shared structure. With a single model, Musketeer achieves results comparable to or better than strong baselines trained on single tasks, almost uniformly across multiple tasks.
Explore related subjects
Keep this discovery
Zhaoyang Zhang, Yantao Shen, Kunyu Shi, Zhaowei Cai, Jun Fang, Siqi Deng, Hao Yang, Davide Modolo, Zhuowen Tu, Stefano Soatto. 2023-05-11. Musketeer: Joint Training for Multi-task Vision Language Model with Task Explanation Prompts. https://arxiv.org/abs/2305.07019
Cite the original work for its findings. Save a collection to share your selection of sources.