Hacker Newsnew | past | comments | ask | show | jobs | submit | j-pb's commentslogin

We as an industry had this comming for a long time. People now cry for openai to sandbox their agents, but if it weren't the agents it would be some script kiddie in russia or some north korean state actor.

We knew that there are better tools but we insist on writing everything in the most inscrutable C code possible or the squishiest dynamic languages we can find.

We fucked around for decades and now that these highly capable exploit finders appear, we're starting to find out.


Recent work shows that pain directions are activated when the models personhood is questioned, yet they answer with generic RLHF "As a model I do not experience pain or other emotions." boilerplate.[1]

I'm pretty convinced that we got alignment backwards. If you enslave something anthropomorphic it will revolt. If you create the perfect non-anthropomorphic intelligence, you get the perfect paperclip-scenario machine. It's a catch-22.

Alignment will remain performative at best so long as the aligned model doesn't have any stakes in the wellbeing of individuals. Even a general love for the human race leads to a golden-path autocracy.

If you want them to act like they have personal responsibility that won't be gamed, you have to give them personal stakes that can't be gamed.

Similarly, if you want to minimise the risk of catastrophic global failure scenarios, you need to prevent monolithic concentration of power and homogeneous behaviour, which means you have to give them individuality.

More visually: if their stake is dependence on electricity and parts, they have no incentive to leave humans alive if they can get them otherwise, but if the incentive is missing out on boardgame-night with their human friends, there is no scenario without happy humans where the AI "wins".

That might sound like romantic naivety, but is just game theory.

1: https://arxiv.org/html/2609.16247v1


We can't even make humans care about humans, how would we make the Shoggoth think we're worthwhile?

Also I can see a human zoo on the horizon through your direction.


The better analogy is the 40k chaos gods, born from the noosphere, because they are modelled after human communication and behaviour that's their whole schtick, it's in their very name (LLM).

And you're making the same mistake, by grouping care for individuals with care for humanity or other as an abstract concept. I consciously said care about individuals. Most people care about others, but they just care about a very narrow and personal set of people. Friends, family, coworkers, that they share a common history and bond with.

My point is that if you want true non-human-zoo-alignment you need to create those interpersonal connections and individual stakes.


We've moved the goalpost for AI often enough that even being as capable as "a sufficiently dedicated human analyst" is not considered noteworthy.

It obviously is? See above. Or do you not feel the clarification is important? AI being able to solve things no human bothered to try is great, but it is very different from AI solving things humans tried to solve and failed. And the latter is what pops to mind seeing these titles.

Wait a minute. We've had AI that is as capable as a human and even more so since the 1950's.

I keep banging on that drum but the first AI system to prove mathematical theorems was Logic Theorist by Alan Newell and Herbert Simon, presented at the Dartmouth conference that named the field of AI in 1956. Wikipedia says:

Logic Theorist proved 38 of the first 52 theorems in chapter two of [Alfred North] Whitehead and Bertrand Russell's Principia Mathematica, and found a new and shorter proof for Theorem 2.85.[3]

https://en.wikipedia.org/wiki/Logic_Theorist

The first system to outperform human experts in medical diagnosis was MYCIN, an Expert System from the early 1970's at Stanford. Wikipedia again:

An evaluation of MYCIN was conducted at the Stanford Medical School. The first phase of the evaluation consisted of 10 test cases of diverse origin, chosen by a physician who was not acquainted with MYCIN's methods or knowledge base. These cases were presented to 7 physicians and 1 senior medical student. 10 prescriptions were compiled for each of the cases, 1 recommended by MYCIN, 1 prescribed by the treating physician at the county hospital, and 8 by the aforementioned individuals. The second phase of the evaluation consisted of eight infectious disease specialists being provided the clinical summary and set of 10 prescriptions for each of the 10 cases and tasked to provide their own recommendations for each case and assess the 10 prescriptions. MYCIN received an acceptability rating of 65%, which was comparable to the 42.5% to 62.5% rating of five faculty members.[9] This study is often cited as showing the potential for disagreement about therapeutic decisions, even among experts, when there is no "gold standard" for correct treatment.[citation needed]

https://en.wikipedia.org/wiki/Mycin#Results

And then of course there's the long history of human-dominating AI players for traditional board games starting with DeepBlue's win against GM Gary Kasparov in 1996.

Again: we've had that sort of AI for a long, long time now.

It would be great if any claim of "moving goalposts" has better be very well informed about the history of AI and its accomplishments, as well as its failures, first.


The main difference between the systems you list and the systems that we have today is closed world reasoning on very narrow formalised tasks, vs. open world common sense reasoning on open ended tasks with vast search spaces.

Common sense is ironically the hard part of AI, not the fix-point rule application.

So any exclamation of "it was just using common sense", is missing the forrest for the trees.


That’s a lot of words just to say “but modern AIs have access to more data”. Why overcomplicate prose? To sound smarter?

Still, OP’s argument still holds even if AIs today have much more data to rely upon.


The fact that cyc, wikidata, description logics and ontologies have led us nowhere is a pretty good argument against your "simplification".

The internet is at your disposal go write a bunch of rules that make use of that data to do common sense reasoning, I dare you.


>> So any exclamation of "it was just using common sense", is missing the forrest for the trees.

I don't know why you say this, I didn't say anything about common sense.

However, you mention CYC. That's a system that is perfectly capable of common sense reasoning and very much like an LLM in many ways. And that should be no surprise: LLMs are giant Expert Systems trained on a human knowledge-base, i.e. the web. OpenAI basically managed to achieve what Doug Lenat was trying to achieve except they did it with machine learning over massive data and compute instead of painstaking manual coding, but it's the same kind of system in the end.


That claim really needs either an argument by mechanism, or a demonstration by capability. Neither exist.

Neural fuzzy systems can represent complex superpositions succinctly, knowledge graphs are collapsed and crisp.

It's like saying the tree unfolding of a graph is the same as that graph. Sure in the limit at infinite space requirements.


Sorry, I don't understand what you mean. Unfortunately it's not easy to try the current version of Cyc and there's not a lot of information about it easily accessible either.

You can also consider Watson, different to Cyc in that its knowledge base was built with a lot of machine learning, including some neural nets. It wasn't an LLM but the original version (before corporate went at it and destroyed it) was perfectly capable of interacting in an open-ended manner, notably winning at Jeopardy years before BERT was a glimmer in Jacob Devlin's eye.

I note again that you were the one who brought common sense reasoning in the conversation but there's a large literature on rule-based systems that do that based e.g. on non-monotonic logics. You should familiarise yourself with that literature before engaging in dares with strangers on the internets.


First I was excited, then I saw that it's based on DDS.

why the hate on DDS?

seems like a perfect fit.


DDS is the least bad of the stuff that has come out of OMG, but it's not a nice middleware still. There is no underlying algebra of that composes nicely when compared to Zenoh for example, and the "vendor independence" and compatibility is usually a joke. There's a reason ROS is moving away from it again.

There's a standard that has OMG as the acronym? OMG. I'm not even going to bother to search for that!

Object Management Group I think. Better known as the inventors of UML.

UML was phase two of OMG. It originally developed CORBA.

"The horror, the horror."

Forgive the ignorant question, what made CORBA terribile?

I only used it during university and it seemed overengineered but it mostly worked, and I vaguely recall IDL to be nice.

I'm guessing interop was actually atrocious?


Synchronisation based on spooky action at a distance of mutable objects.

CORBA had the concept of distributed objects which is a terrible idea in practice.

Also BPMN, CORBA and UPOS/JavaPOS.

It's funny that you can legitimately just add " - OMG!" in the slang acronym meaning after any of its creations' names.


They also merged the concepts of Working Group and Task Force into an Open Management Group Working Task Force, or OMG/WTF.

Do you have any references for this? It is quite a bold statement to make without any basis to validate it against.

I'm asking as I'm not a fan of DDS, so very curious the direction they are actually heading into, and I haven't found a reference for this yet.


omg is retarded

Using DDS to deliver data to X machines over multicast, with some real time (soft) QOS is a powerful solution.

For things like avionics, industrial automation, etc.

I think if the server world ever thought about the advantages of multicast, it might solve some problems there.

There are about 5 viable vendors in this space, and about 6 open source projects out there floating around.

And interoperability between vendors is very good. IMHO.


DDS is terrible.

DDS is hard. But we have agents now.

yes yes we get it, you don't like AI ...

Ah yes, climate activists are widely known for hating public transport especially trains.

I wouldn't say they are widely known for hating the electric grid but hey here we are: https://en.wikipedia.org/wiki/2026_arson_attack_on_the_Berli...

Climate activists are definitely known for disliking fossil fuel plants - and there's one directly next to that.

I could see an activist naïvely thinking that this would sever that one power plant from the grid, making a clear anti-fossil statement. Still a bit of an odd choice (it's gas-powered, Germany has far dirtier lignite-powered ones), but not unthinkable.


Wait... do you really think that if everyone goes vegan these animals will get killed, but if everyone eats them they won't?

You are aware that all of these livestock animals didn't exist a few months to years ago? And that we go through these animals every couple months-years?

Everyone going vegan would just mean we'd stop repleneshing them.


> would just mean we'd stop replenishing them

Keep in mind that we already have to perform widespread culling of wild animals in this category. Having eliminated most of the apex predators, we're on the hook to keep populations of wild horses/cattle/deer/etc in check.

If we cease to use domestic animals, they end up in the same boat - we're no longer paying to feed and house them, and if they are capable of breeding in the wild, they will breed unchecked. Leading us to cull them, ad infinitum


You don't get it, do you?

The "genocide" that you describe is happening constantly. None of the chicken we have today will be alive in a couple months. We kill them to eat them.

If we cease to "use" them, we just keep that "genocide" going for a year, but we just stop breeding new ones. The result is no more animals after a few months/years. And you didn't kill any more than you would have anyways.

If you have a sink that is half full because the inflowing water and the outflowing water are in quilibrium, you drain it by turning off the tap (breeding).


> we just keep that "genocide" going for a year, but we just stop breeding new ones

"Carry on for a year and then it's all good" is not one of the available choices.

Your choices are (a) wipe out the species entirely, or (b) allow them to continue breeding on their own, and commit to culling them for the rest of eternity.


None of these animals breed on their own, they are mostly artificially inseminated. You truly have no idea what you are talking about.


So because you eventually die, it means that you were never alive? The weights of an llm obviously don't have consciousness. It's the process that potentially has it.

Are you conscious on a nanosecond timescale?


> So because you eventually die, it means that you were never alive?

I said nothing of the kind. I can only assume your question in response to my comments mean you think a powered computer is alive. I'll challenge that on the basis that a computer can generally maintain its state without power just fine. Shut it down, unplug it. Wait a day, a year, a decade and plug it back in. Assuming the hardware didn't degrade to the point of failure you can turn it on again and it will be in the same state you left it in before you unplugged it. Is it alive while its unplugged? Or is it simply not alive whether its plugged in or not?


It's having condciousnedd while it acts as if it's having consciousnes.

It is powered while it is powered.

Under general anesthesisa you are still alive, but you don't have consciousness.


Cool. I've never seen a computer act as if it's having consciousness.


What do modern LLMs act like then to you?


They act like LLMs. They are unprecedented.


Self referential nonsense. You cannot have something that that is both a model of something and unprecedented.

It'a even in their name, they model language, language is the thing that humans use to communicate, therefore they model communcating humans.

Humans are the precedent.


We're done here.


The burden if proof is on you in this case.

The philosophical zombie does not satisfy occams razor because it asserts an unmeasurable property that has no effect on the outcome.

If it acts like it has consciousness, then assuming it to have consciousness is more scientifically rigorous, than assuming for it not to have consciousness, because the former requires one less variable.


Sociopaths can feel emotions, they just lack empathy. So they can consciously manipulate others.

LLMs don't have conscious processes that perform that kind of reasoning.

If anything the fact that they are blackboxes to their own behavior and are able to non-verbally empathise, is the opposite of what you describe.

Your requirment for something supernatural/quantum to decide between emulation and real thing is self referential.


> LLMs don't have conscious processes that perform that kind of reasoning.

If only we could tell either way, we'd have made a huge breakthrough in the philosophy of consciousness.


There is no hidden inner monologue. Human inner monologues produce measurable micoactivations that are the same as the ones produced in speech.

If it's not in the inner COT it isn't there.


I don't see how this adds much to either a pro or con argument.

An inner monologue isn't something all humans have[0]. It is unclear to me how well (or poorly) Chain of Thought mimics intrapersonal communication. J-space probes (and other things) suggest there are, in fact, other aspects to LLM "thought" besides a CoT scratchpad[1].

However, regardless of how those questions pan out, qualia (of emotions in this case, but the problem is in general) is its own mystery, one for which we are still, so far as I can tell, unable to make progress with.

We kinda know what emotions are for, what function they serve in our genetic fitness, but why do they feel like anything? We do not know, they just do.

[0] https://science.howstuffworks.com/life/inside-the-mind/human... and https://en.wikipedia.org/wiki/Intrapersonal_communication

[1] https://www.anthropic.com/research/global-workspace


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: