hello, in the inference.py you offered in #14, I see the multi-modal input tokens for LLM, it includes bbox token, but I can't find where you replace the bbox token or you use the image feature which got from clip and interpolate. Can you explain it for me? thank you.
hello, in the inference.py you offered in #14, I see the multi-modal input tokens for LLM, it includes bbox token, but I can't find where you replace the bbox token or you use the image feature which got from clip and interpolate. Can you explain it for me? thank you.