Hacker Newsnew | past | comments | ask | show | jobs | submit | jpadkins's commentslogin

It was part of Sundar's Q3 OKRs. He is going up for promo this cycle.

Is he running for office in November?

This is what I did. Hope it works out. The other benefit is you have a more natural method to avoid lock in. A lot of "improvements" to the agent harness I believe are attempts to build customer lock in.

Questions for those who insist on reviewing all code: Why do you not the output of the compiler? Do those reasons not apply to agentic output?

I think a group of us have found, for certain domains, the agentic output is passing all of our standards we have for correctness (ad hoc tests, automated tests, etc) so we don't review the code, just like we don't review the assembly or bytecode. The compiler and assembler have bugs, and its more likely the agent system has a higher probability of bugs. But it's still good enough for personal projects, web apps, prototypes, internal tools, analytical helper tools, etc.


Engineers who work on high performance systems actually do review compiler output.

But compilers are deterministic and have close to 100% test coverage. Bugs in them are rare and hard to find. None of this is true of AI.


AI does not create a new layer of abstraction; a compiler does. This is computer science 101 and I'm kind of shocked that so many engineers make this basic mistake.

How come AI is not a new layer of abstraction?

AI translates prompts in natural language into code in a programming language, just as a compiler translates a programming language into assembly.


High-level programming languages use well-defined, limited syntax to encode program behavior in a deterministic way. Compilers translate these directions into programs that should have identical behavior between compiler runs and between different compilers. (If they don't, that's a compiler bug, or else an artifact of UB.) This means you can structurally and provably rely on the high-level programming layer without worrying about what's underneath: the assembly is abstracted away. You also get the massive benefit of being able to tweak and improve the assembly layer independently of the programming language layer, so long as the contract between language and compiler is not violated.

If AI is a "layer of abstraction," so is working with a bunch of contractors to make your app. The definition is broadened to the point of absurdity. (Solely: "I don't have to think about code anymore," which only touches on the most superficial aspect of an abstraction.)


Why can't an abstraction be unreliable and nondeterministic, maybe it's just a bad and leaky one. LLMs are algorithms, i.e. Turing machines after all. They can often produce a program without you even looking at the code, hence they abstract coding away.

Honestly even calling a contractor a layer of abstraction doesn't sound too absurd to me.


The output of a compiler is predictable and deterministic. LLM output isn't.

Correct, but for some projects predictable and deterministic aren't required (or sometimes even desired)... Think creative pursuits.

I kinda agree, maybe, but I don't think compiling code is a creative pursuit.

you didn't really answer the question. You just stated the idea is silly. Reasoning is not in the same class as soul, heaven or hell. It's not obvious that human reasoning is not related to an inner monologue. And that LLM chain of thought process is approximating inner monologues.

Our prefrontal cortex are signal prediction 'machines' so when a system that has a signal prediction core has attributes that are similar to our brains, we shouldn't dismiss it out of hand.

I find people that take this line of argument attribute too much supernatural or magical properties to our own brain and nervous system.


Hmmm, ok - let me answer the question clearly then. LLMs cannot reason, because following their own instructions and algorithms is not reasoning.

Humans can't reason because following their own electrical and chemical instructions and algorithms is not reasoning.

I have found that having a separate agent (session / instance) do the benchmarking and reporting the results back to looping optimizing agent is a clean way to prevent cheating. The benchmarking agent has no reason to cheat, its goal is to just to run benchmarks when tickled.

I also found this is really nice for quality evals. Have one agent with no context on how something is made do a quality review, with lots of detailed feedback. Then pass back the review notes to the implementor for feedback. It works a lot better than having an agent self-evaluate its own quality.


Has there ever been an instance in history when this strategy worked? Plato argued that writing things down will make your memory worse, and less skilled as a debater (kind of true!) How are the Luddites doing at textiles? I remember the arguments that using 'high level languages' like C and Pascal will make you not understand machine specific details (kind of true!)

I respect that you want to learn how things are done, that is a great trait. But once you learn how its done, you should use the tools to free up cognitive load for more difficult tasks.


> There are ample reasons to believe that Elon Musk is running his mother's account

Yeah, I'm sure the guy has time to run his mom's social media account. What's your reasons or evidence? Did you consider maybe it's a social media manager one of them hired running the account?


what are the usual reasons?


See later, i.e.:

> somehow this high profile AI model seems more disgusting than others and it is in a way impressive.


why? Grok means to understand intuitively. Seems like a pretty good name compared to Claude or ChatGPT, no?


It's intuitive if you're a nerd, maybe


who the fuck is using reddit for product recommendations? reddit was astroturfed in like 2015.


If not earlier. Reddit has been highly manipulated since at least 2015.

I used to buy/sell reddit usernames for astroturfing and it was always hilarious seeing one of my usernames on the front page.


Consider applying for YC's Winter 2027 batch! Applications are open till November 2.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: