Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

VLM != LLM. Vision language models basically treat text tokens and image tokens the same. Post-training an LLM on images+text can improve its capabilities. Id recommend searching around the keyword VLM to find more resources on how multi-modal AI works.

- https://huggingface.co/blog/vlms

- https://en.wikipedia.org/wiki/Multimodal_learning

 help



Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: