open-source multimodal large language model that can hear, talk while thinking. Featuring real-time end-to-end speech input and streaming audio output conversational capabilities.
thanks for you contribution, To avoid other potential impacts of directly exporting PYTHONPATH, I have documented similar issues in the 'common issues' section of the README for reference.
Avoids #25