Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

When an LLM can re-tokenize on the fly, due to newly learned data, let me know.

It won't prove intelligence, but at least it won't be static like a book.



https://x.com/sukjun_hwang/status/1943703574908723674

> Tokenization has been the final barrier to truly end-to-end language models.

> We developed the H-Net: a hierarchical network that replaces tokenization with a dynamic chunking process directly inside the model, automatically discovering and operating over meaningful units of data


Good to see. Will be better to see once the scope is clear.

But something such as this is required to move towards actual intelligence.


This is a dumb critique. A thin wrapper running a new samples of training data and updating weights is something already done in many situations. A sophisticated rag system incorporates new information even if the weights themselves aren't updated, effectively giving "new memory".

LLMs have problems, in practice being "static" aint one of them.




Consider applying for YC's Winter 2027 batch! Applications are open till November 2.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: