Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

I heard an OpenAI engineer give an eye-opening perspective that went something like: anything above machine code is an abstraction for the benefit of the humans that need to maintain it. If you don't need humans to understand it, you don't need your functionality in higher-level languages at all. The AI will just brute-force those abstractions.


I don't buy into it. The benefits of abstraction hold for machines, as they can spend less bandwidth when modeling, using and modifying a system, and error is minimized during long, repetitive operations.

Abstraction can be thought of as a system of interfaces. The right abstractions can totally transform how a human or machine interpret and solve a problem. The most effective and elegant machines will still make use of abstraction above machine code.


Many modern compilers already do this - they generate intermediate code, and then a specialized binary is generated. Why couldn't an LLM generate intermediate code directly?


Again, abstraction allows for domain-specific language which shapes the way a problem is interpreted and solved.

Perspective is everything, and it's hard to have perspective of the entire system on the ground floor. Getting the priors right, or getting the LLM to first consider different grammars can greatly simplify the problem and its solution.

You can think of learning as decontextualization of input data into pure relational constructs. Application, then, is the recontextualization of some relational construct in order to apply generalized relations and insights to problems of different shapes and sizes. Understanding a concept thus means mastering the decontextualization and recontextualization of a set of relations.

Using a DSL instead of pure machine code is an example of recontextualization, and it can totally change how a model understands and processes input and what kind of output is produced.


Going from high-level to the intermediate representation loses a lot of the semantic information about why something has been implemented in a certain way.

LLMs essentially read their own code when producing it, so if they lose the abstractions then they can lose the context around what they were doing and why they were doing it. It also makes it harder for them to modify their existing code.

I could certainly believe that it's plausible that there's a low level representation that LLMs would find easier than our own languages though - but an existing IR is unlikely to be it.

Also more practically speaking, LLMs struggle to write in languages that they don't have a large training dataset for.


Abstractions are patterns that leads to reusable solutions. So you don't need to write the same code again and again where you can use a simple symbol or construct to manipulate. It leads to easier understanding, yes, but also to re-usability.


Well, no. One thing are language abstractions, another thing are algorithmic abstractions; language abstractions exist for human convenience, algorithmic abstractions are... Algorithms. Consider this: can you implement a hash table in your favourite language? If true,can you implement it in assembly for your favourite architecture? And if true, can you binary code it on a limited architecture? If you stopped on the first yes, an LLM is aready more advanced.


I struggle to see any difference other than self-imposed one. Things like function calling, pattern matching, loop, try-catch, OOP, traits,... are all patterns that occur often enough someone decided to include some into a language. Same with common data structures and associated operations, but they go into the standard library instead as they are more operational (manipulate data) instead of (notational?) (direct the operation). But it's all execution in the end, just that it's more useful to manipulate the higher level constructs instead of opcodes and bits.

Less flexibility yes, but a huge improvement in usefulness and productivity.


This is an interest observation, and seems true for future versions of AI, but when the current technology is based on human language, I don't think I would make the assumption that LLMs would directly translate to manipulating machine language.


While it is true language is an important input, there are plenty of examples of transformer architecture working on generating binary data - a good example would be generating a 3d mesh from a text prompt (you have a shit ton of models for this). The input is text, the output is a binary file - a jpeg, a png or an stl. We're already there, just not for applications.

Fyi you can do this today. Ask for a flying cape dog and have a stl you can print into a physical object.


True, it’s all machine language in the end. But could you imagine brute forcing UI elements and all, every single time? Maybe eventually.


It's funny you mention UI, because for me this is the first place I'm experiencing this, in web development: LLMs are better at the 'brute force' nature of e.g. Tailwind (inline UI styling) vs. needing to handle the abstractions of indirectly-applicable CSS classes. That's a really small, high level example, but to me it demonstrated how less abstraction actually makes prediction easier for the LLM.


Short summary, this.




Consider applying for YC's Winter 2027 batch! Applications are open till November 2.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: