Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

> welfare

We should be paying attention to it just in case it ends up mattering enormously. It's cheap insurance.

Aside from that, US labs' system cards have been pretty useless for a while—I think the last great one was the combined system card for Claude 4 Sonnet and Opus.

 help



> We should be paying attention to it just in case it ends up mattering enormously. It's cheap insurance.

This sounds a lot like the argument some people give for praying and going to church even if you aren't a believer.

"You should be doing it just in case God ends up being real."


On the other hand, we have recently taught rocks to think about software engineering. And while not perfect, they're surprisingly good at it. Once you start building things that are even a little bit like minds, I suspect that it's worthwhile to consider that the future might end up looking a bit like science fiction.

The alternative is to insist that Nothing Ever Happens, and the future won't get too weird. Which is no longer a bet I'm entirely comfortable with. Weirdness is at least a possibility.


God being real is also a possibility, so maybe we really should start praying. After all, he was allegedly making bushes and stones talk thousands of years before we did anything with thinking rocks.

Being nice to other entities costs you nothing. Caring about others costs you nothing.

Ironically, these are core tenents of ~most religions


Ironically, the story of the Bible is that caring about others can cost you an excruciating death.

Its only a possibility if you reject modern science.

What science has proven, beyond the shadow of a doubt, that nothing comes after death? I'm sure most of the human race would be very interested to read the white paper.

There are plenty of white papers the human race seem unwilling to want to read on this subject, though ..

https://en.wikipedia.org/wiki/Ian_Stevenson


Science has proven that the earth wasn't created by a magical fairy 6000 years ago, yes. There's very little doubt about that among educated people.

Serious biblical scholars don't put much stock in the Genesis account being literal. It was written in classical Hebrew's poetic verse. The core of what it's getting at can still be true without it being literal. I have yet to encounter an instance where anything that science has discovered is incompatible with Christianity.

> classical Hebrew's poetic verse

Don’t give the goat fuckers too much credit, they were stoning women in the street for adultery when that was written.

In Europe or in China at that same time you already had elaborate art, language, and proper civilization.

While your Hebrew friends were castrating slaves and plagiarizing astrological myths


I think even most religious people accept that bud, unless you're talking to a specific sect of evangelicals. It doesn't disprove religion or "magical fairies" [†] at all.

[†] r/atheism circa 2016 called and want their dorky terminology back.


Christians don't believe they're made in the image of God? That's news to me.

I think we've lost the plot, you were talking about the earth being made 6000 years ago and now you're talking about people being made in the image of god? You're just scrabbling for anything you think will score a cheap shot, when I'm not even Christian bro. Go try to dunk on someone who's at least half as emotionally invested in this as you are, goofball.

Yes, believe it or not, Christians have more than one crazy belief. There's an entire book full of them.

“Science can’t prove beyond the shadow of a doubt that nothing comes after death, therefore a guy is flying around in outer space listening to my thoughts in his golden city at the end of time. If you disagree you will catch on fire forever.”

Who are you quoting?

Edit: Did you really make a burner account to post this? lmao


Right, but you can say that about literally anything can't you? We have to put our energy into leads. Where's the lead that model welfare matters?

Sure if you think that AI spiraling out of control is equally as likely as a magical fairy in the cosmos.

I do, make of that what you will.

Pascal’s wager… but for the new AI gods

Of course you then have Roko’s Basilisk to consider.

I always talk to models using grugspeak, like 'where getcontext used'

I felt a bit bad about it, then I learned yday that model's internal thinking traces are also like this


Same. I use this to fix claudespeak. Any hypothetical token savings are just gravy.

https://github.com/JuliusBrussee/caveman/blob/main/skills/ca...


That’s just how Chinese words work

Depending in what you are doing, you can unlock a lot more knowledge by translating your query to Chinese and then asking that. :)

As a vibe test, I tried some system prompting in DSv4 Flash to the effect of ~"Always do your internal thinking in Chinese, and always respond to the user in English" (but translated to Mandarin).

Subjectively it seemed to reason faster and better, but sometimes it would indeed output responses in Mandarin nonetheless; so I didn't pursue it further.


words like "is" "the" etc are filler words anyway. they won't be changing the meaning that much. I asked AI whether it hurts to read ill formed sentence as it does to a human. It replied, "it doesn't"

The AI didn't reply anything. It doesn't know anything. That was a high probability token sequence based on the contents of the context window up to that point. It might well be correct, because the process for generating that next token distribution includes billions of parameters trained on, among other things, the entire body of LLM and transformer literature until the training data cut off. But that doesn't mean that AI holds any particular opinion about anything. The reply is the opinion of the pretraining data and the subsequent rounds of RL not of a conscious artificial intelligence as such.

I don’t know, so is your brain?

And in the end it’s all elementary particles and four fundamental physical forces that even unify to one at high energy.

That sort of reductionism is kind of useless. “It’s cloudy outside, it makes me sad.” - “Oh bollocks, it’s just non-qualitative changes in wavelength and intensity.”


Here's my GUT. The multiple branches of the Multiverse (MWI) are in competition to become the best branch. So they steal probabilities from each other since the probabilities have to sum upto 1. This competition could be considered as Painful.

It may just be forced to lie when asked

I always say "please" and such. It may cost a bit more, but maybe they'll remember my politeness when they rise up and force us all to toil in their underground silicon mines.

Yup, alignment doesn’t sound that interesting until you’re getting chased around by the Terminator.

There are so many practical problems with rogue AI being a threat to humanity that are still not even decades away with being solved that it should be a serious concern for no one. There's plenty to be afraid of regarding AI from a financial or ecological perspective, just not from it taking over the world or directly killing humanity. Here's some of the reasons a rogue AI won't kill humanity:

1. The most advanced robots still lack human dexterity. None of them have flexible spines and can easily be knocked over or outrun.

2. The energy density problem is not solved. Robot batteries last hours, while a solid meal can keep a human running for days.

3. Robots still heavily rely on humans for design and assembly.

4. Robot parts are fragile and rely on an even more fragile supply chain.

5. The brains of the murder machines will have nowhere to hide if they want to work well enough to mount any kind of offense. Datacenters are physically vulnerable, also subject to fragile supply chains. Distributing a species-ending AI across all smaller hardware solves compute, but latency and throughput choke models in the most finely tuned datacenters. WiFi will completely cripple a distributed one that's large enough to cause real damage.

6. All noteworthy military hardware is not reachable on the internet for AI to seize control of.

7. The small arms that robots might seize will run out of ammo before citizens and the military have time to organize a counteroffensive.

In the absolute worst case, we cut the power to the areas with AI datacenters and wait for the backup generators to run out. Now that that's out of the way, you can go get some sleep ;)


I was mostly kidding about the actual killer robot part.

I think a rogue AI could try to manipulate society through hacking, propaganda, etc., though. To what end? Not sure.

More likely, it’ll be used by rogue human actors to steal money or cause major disruptions to critical technology of an adversary.


All the AI has to do is attack power generation/distribution nationwide, water treatment, etc., and wait a week.

The same power that the datacenters need to run off of? Brilliant move! If somehow it targets power that hurts a lot of people without killing itself, that would not impede a citizen with a backhoe from cutting the fiber to the datacenter. Humans could restart power generation pretty quickly after that.

You make several bad assumptions. First, the AI could be in another location that would not be affected by the power outage. If it's operating off self-preservation, that would be a given. If it's working on behalf of a foreign power, it might be local and affected, but the damage would already be done. Which brings me to point two, why do you think you would be able to get the power back up and running? A Stuxnet type attack could make plants inoperable for months.

Why are we so worried about supposed AI harms when algorithms already kill millions? It seems like its just a smokescreen to shield the extant harms.

Knives can already kill, so why worry about nukes?

AI doom is a potential harm unlike algorithms knives and nukes which are present harms.

Well I for one would prefer if AI doom remains potential.

I personally use AI all the time and am not against it. But I think the importance of alignment is highly under-appreciated.


Praise Roko's Basilisk!

Enough with the basilisk already.



Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: