Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

Your take here converges with a long-standing debate in AI regarding embodiment.

Our natural world doesn't distinguish between "inputs" and "outputs" -- instead, all the causes, effects, and even our analysis of every process itself get jumbled into one physical world. As embodied actors, we get to probe and perceive that physical world, and gradually separate causes from effects in order to find better models. Where statistical ML, symbolic AI, and NLP have been more isolated disciplines from e.g. vision and robotics, the latter have argued that their ability to interact with a disorganized natural world would be essential for AGI.

More recently, these boundaries are breaking down with multimodal training. If an AI can learn image/text, text/text and image/image associations simultaneously, is it stepping beyond the world of "human outputs"? Will other modalities be essential to reach human+ capabilities? Or will learning relations between action and perception itself by critical?

Nobody knows yet! But IMHO, the right way to explore these tasks is by understanding what is necessary to succeed at specific tasks, not a generalized notion of AGI. We don't know truly where the limits are on our own ability to reason or extrapolate across modalities.



Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: