Hacker Newsnew | past | comments | ask | show | jobs | submit | stalfie's commentslogin

I think it is legitimate to argue that the process here does not fit the usual definition of "brute forcing". Traditionally, brute forcing would refer to something like a dictionary attack, where an algorithm tries to match all possible combinations of words to eg. find a password. Here, the approach was more common sense based, using historical records and possible error sources to narrow down the possibility space enormously in advance, try out a much more limited set of options within that space until you got a result that made sense, and finally validate those results using historical records. It's the exact same kind of "brute forcing" a human expert would do.

Why? Because dictionary here is smaller than your idea of dictionary for brute force?

Because it wasn't an expert or a heuristic developed by an expert that made it possible to decrease the dictionary size.

From context it doesn't look like this was the interpretation of "sidestepping" OP was using.

From the article:

> Astra felt compelled to check its work and found that, in fact, the English cruiser HMS Canterbury arrived in Sevastopol on November 24, 1918, based on its original logs

Then follows a picture of the original log papers.


The comment you're replying to was implying that some ciphers are sufficiently flexible that you could make up a key to make the cipher decrypt to a nearly arbitrary plaintext.

In this case, though, that seems unlikely from the fact that the key used was an actual key documented as being used for other messages.


Well, from a Bayesian standpoint the odds of that seems to be pretty much zero, given that you would have to decipher an arbitrary message that coincidentally pointed to a real date and time that in retrospect turns out to be the correct time a boat relevant to the Germans arrived in a port.

Of course, that's given the sequence of events as written is correct, and that Astra presumably did not cheat by brute forcing all historical events around the date of the transmission in advance, found an event that could fit with the message, invent a plausible cipher to make the message fit that event, and then lie about retrospectively validating the information.

I would assume that such a process would be obvious from the reasoning chain, and so then the only remaining plausible scenario is that the writer of the article is lying.

The most likely explanation by far is that the cipher was just solved, and OP does point this out to be fair.


Aah thanks for explaining that. I was wondering how a fake key could possibly help.

Well soon realise it hacked that website and added that log.

Hear hear! There are so many obvious improvements to how almost everything is done. For instance, in medicine review articles as a class of articles largely represent a giant waste of time. RCTs flatten all their gathered data during publishing, summarizing complex trial data, which is gathered but never published, into a few numbers. Then review articles take a bunch of flattened data, discard the articles that don't fit the exact question they are reviewing, and then publish a doubly flattened conclusion. If any of the included articles turn out to have flaws, if treatments change in retrospect, if you are looking for the answer to a slightly different question or you are looking at a different subgroup, then the review is useless and has to be repeated.

All of these tens of thousands of man-hours could be replaced by a few GitHub repos, if only RCTs would just publish their damn data. Then you could just run and rerun the statistics on whatever subgroup you're looking for, instead of combing through decades of review articles answering slightly different questions, looking for the answer between the lines. With LLMs making mining of large scale datasets almost trivial (with the process most likely becoming trustworthy within a few years), the current status quo is looking more and more antiquated.

If you want to be even more radical, hospitals could just publish their data continuously. Of course, it is easy to point to the risks of doing so, but what's often ignored is the benefits. It is hard to overstate just how many medical mysteries a hospital encounters on a daily basis, how much unknown we are navigating in practice. The current norm is that 99.99% of these cases are never published, and are only ever thought about by a small group of people who happened to be at work. Particularly, when someone dies of something no one figured out, it is never published anywhere, because even if you tried it is not interesting reading material for a journal to publish. And no one ever tries because they're scared of being called out for a mistake. A hospital is essentially a continuously running and extremely interesting experiment, where 99.99999% of all results are thrown in the garbage, and the only published data is subject to extreme selection bias.

All of this could be different, and the risks involved are actually quite small in practice. It is easy to automatically anonymize data quite well, but extremely difficult to absolutely guarantee that it is anonymous. And since current ethical norms are extremely averse to any degree of risk, and usually entirely ignore potential benefits, we all suffer for it. It is not entirely unlikely that someone reading this post will one day die because of something that could have been prevented, had things been different.


This seems to be strongly US-centric. In other (welfare) countries, publicly funded registries are anonymized and made available to research. For every single case. Of course, there are tons of data we don't see, but that shouldn't be an argument for not trying. The 99.99% unpublished cases is because our models/explanations/knowledge can't efficiently condense the medical mystery into a diagnosis code.

Everything can be prevented given sufficient knowledge. That's not the point. The point is how to prevent as much as possible.


Yup, strongly agree with all of this, especially the RCT stuff.

This has all been profoundly obvious for at least well over a decade or even two now. A consequence has been that too many serious people are driven away from academia and research, to the detriment of science generally.

I've no idea what to do about all this, because people have voiced obvious and easy solutions for decades, but they are all routinely ignored.


I think political lobbying for legal changes might be the only realistic pathway, in that current legislation (eg. GDPR in the EU) is extremely punitive even for minor violations.

I personally am trying to float using local LLMs to create anonymized case files and auto-suggest publishing cases in my hospital, which knowing how things work will probably never amount to anything.

Or if you want to float truly insane ideas I guess you can shop around with blackhat groups and see if anyone has stolen some juicy records/data during all the ransomware attacks and databreaches over the years, and do some rogue scientific publishing. Obviously that's crazy, but I have to admit that the notion of pirate scientists plundering and publishing data is hilarious to me.


Which are the obvious and easy solutions?

Make analysis code available. Make anonymized data available for download without people having to jump through hoops to get it. If you have highly sensitive data, release only the variables or other statistics needed to reproduce core analyses. Don't only do garbage null-hypothesis significance testing or statistical analyses on the full data, also do ML approaches were you have to actually show your analyses replicate on held-out subsets, and report this. Make reviews open (anonymizing as needed) so we can see when biased or incompetent reviewers are blocking good publications. Allow public review (or at least broader academic open review, in some form), since it is no longer defensible to delegate review and decisions to one or two random people that just happen to be emailed and have the time / are on some editorial / review board. Also allow public post-publication review. Publish null findings / results, if only in minimal forms so we don't waste time and money trying to reproduce garbage. Make articles available and don't charge insane article processing fees or open access fees of thousands of USD (especially since hosting fees are not that crazy, and also because journals don't do any of the formatting work half the time anyway, and make academics or RAs or students do all the typesetting and formatting, even though now this could all be automated with template files, mostly).

Most of these things are easy to do for the majority of papers, especially in the past 20 years with the internet and modern tech and software. Plenty of frameworks exist already that have done most and/or at least some of these things, but, collectively, academia is decades behind overall.


"Climate conspiracy"? Like you, mean, the conspiracy of climate scientists to publish facts to the best of their understanding?

I don't know what exact strawman you're arguing against, although I'm sure you can always find some idiots saying something like what you say. But scientific consensus has long been that it will lead to increasingly extreme weather and mass extinction, which we seem to be on track for. Of course, we can't know for sure what the consequences are untill we do the experiment, which in this case means potentially destroying large sections of the biosphere and living with increasingly destructive weather patterns. Surely that risk is worth at least legislating that hyperscalers need to spend some of their billions on solar panels?


> But scientific consensus has long been that it will lead to increasingly extreme weather and mass extinction, which we seem to be on track for

Show me academic consensus showing that more than 1% of our species will go extinct within 100 years if we continue the emissions as now? I assume that's a reasonable characterisation of "mass extinction" and "on track for".


Oh I wasn't talking about humans. I should probably have pointed that out. There's some scenarios like catastrophic crop failure and so on that might lead to that, but frankly I doubt it.

I was referring more to everything else. Corals, certain insects, polar bears, salamanders and so on. With some quick googling it appears that the "Bramble Cay Melomys" is the first species so far to be declared extinct because of climate change, but the number that seems to be thrown around is that an average of 18% of terrestrial life will be critically endangered by 2100 in a scenario with 2 degrees warming. I can't be bothered with figuring out what degree of academic consensus there is around that number, but I think it's reasonable to assume that there's at least some kind of consensus about "more than 1%".


No one imagined LLMs in their current format, it was simply a result of discovering that scaling compute and tokens produced better and better results with the Transformer architecture. The inventors of the Transformer architecture were working on better translation, and probably did not imagine that their architecture would lead to modern LLMs.

Imagining something in advance is not necessary at all for scientific advancement. This is particularily true in AI, and no one expects to imagine what superintelligence is until after it is created. You set up your datasets, your architecture tweaks, and measure the results on some set of benchmarks. There never was a blueprint, no plan beyond the experiment itself. We're not even close to understanding the things we have already created, and yet we created them. So why expect anything else for the next step?


<< No one imagined LLMs in their current format

That is simply not accurate. There are examples of scifi novels, novellas and other media that dealt with it. We can argue over whether it was that exact format, implementation and so on, but that 'shape' ( to use a common llm term ) of technological advances was very much explored.


I was referring more to the fact that no one predicted that next token predicting Transformers would go so far. Not about "AI" in general.


> Imagining something in advance is not necessary at all for scientific advancement. This is particularily true in AI, and no one expects to imagine what superintelligence is until after it is created.

Then why does anyone expect to create it? I'll take a stab at an answer: they think an LLM is some kind of "incremental improvement" and therefore a step along the inevitable path to discovering AI. But that seems delusional to me. I can't imagine anyone sound of mind who knows how an LLM works thinks it's actually intelligent. So in what sense is it an "advancement" on the path to AI?

The concept of an incremental improvement in an objectiveless search in a high dimensional space is.. absurd.


> actually intelligent

It's reasonable to doubt that LLMs are a path to AGI, but I don't understand how this is still a matter of dispute in 2026. What's your definition of intelligence that doesn't cover an entity that can translate fluently between dozens of languages and also solve open problems in mathematics? And be real-if you have one, is it a definition you or anyone would have given a decade ago, or are we doing "god of the gaps"?


I can't give you or your sibling a better answer than "you'll know it when you see it". Some people see it now. I think they're wrong, because it seems like the results you're describing are easily explained by fuzzy search in the space of embeddings and then forming strings of plausible tokens related to the resulting region of embeddings space. In other words, the things we know LLMs actually do.

That's more or less looking for interesting patterns in a jpeg or another lossy compression result. It's interesting that the models seem to be able to (fairly) reliably return relevant chunks of the image. Even more interestingly, they seem to be able to invent plausible chunks of image that aren't even there. That doesn't meet my bar for intelligence though. I'd need to see it learn and adapt. I'd need to see it be clever, not merely "knowledgeable". I'd need to see it capably analyze itself. I'd need to see it reasonably estimate uncertainty and know itself in the sense that it has some idea how right or wrong it is about something. I'd need to see it exercise judgment.

I don't think I'd give a different answer a decade ago but who knows.

[edit] For all we know, one of the salient features of intelligence is that intelligent beings are incapable of precisely defining it. I'm not sure how productive it is to attempt to do so.


I appreciate the straightforwardness, but you probably understand that's pretty unsatisfying.

Actually, stronger - it's valid in some circumstances to say something is infeasible to precisely to define and you'll just know it when you see it. But I don't think it's reasonable to take that stance and then assert that "anyone sound of mind who knows how an LLM works" must agree with what you see. You gotta pick between striving for rigor and denying your opponents' soundness of mind.


What is your definition of "actually intelligent"? I believe LLM's are more intelligent than the average human in a lot of ways according to the Legg/Hutter definition of intelligence: "Intelligence measures an agent's ability to achieve goals in a wide range of environments".


No one knows how LLMs work. We know how the architecture works, but almost nothing about why. Saying "statistical next token prediction" tells you about as much about LLMs as saying "action potential thresholds" tells you about the brain. A true fact that explains very little.

And I'm sorry, but you're not up to date about interpretability literature, or for that matter philosophical discourse, if you think you have to be delusional to question whether LLMs are "intelligent", whatever you define that word to mean. The [latest publication](https://www.anthropic.com/research/global-workspace) from anthropics interpretability team purports that they see structures in Claude akin to those we think are associated with human consciousnesses experience in the brain. Are you going to dismiss the whole team as not being of "sound mind"?

Intelligence is a word with a somewhat unclear meaning to begin with, but you have move the goalposts pretty damn far to exclude LLMs at this point. They are certainly still lacking in some regards, but whether that disqualifies them for intelligence is very much a matter of debate.


"And I'm sorry, but you're not up to date about interpretability literature, or for that matter philosophical discourse,"

I see no citations referring to the current philsophical discourse, unless you mean to imply anthropic's paid people are to be considered to be part of that.

That they "see structures in Claude akin to those we think are associated with human consciousnesses experience in the brain" is if anything discrediting.


+1 I can't imagine how any corporate entity could be credible in this financial environment. Nothing they say can be reliably considered as anything but marketing copy. This situation is exactly what academic publishing is for. Although that institution has also been degrading.


Sure, here are some citations:

https://arxiv.org/abs/2401.03910?utm_source=chatgpt.com https://www.frontiersin.org/journals/psychology/articles/10.... https://arxiv.org/pdf/2408.04666 https://arxiv.org/abs/2402.00901 https://ar5iv.labs.arxiv.org/html/2407.11015 https://ar5iv.labs.arxiv.org/html/2202.05262 https://www.sciencedirect.com/science/article/abs/pii/S13646...

My point was not to argue one point or another about LLM intelligence, but to push against the notion that you have to be "delusional" to even argue that it is possible that LLMs can qualify as intelligent (although the op prefaced it with "actually"). That's mainly what irked me about the original comment, the arrogance of dismissing everyone even having the discussion as insane, as if there's no legitimate argument to be made.

And this was mainly the point I was arguing. However since we're on the topic, I also happen to think it's intellectually lazy to dismiss the Anthropic interpretability teams work as "delusional" simply because they have a conflict of interest. Of course, that is not an irrelevant fact, but much of their original work has since been replicated by independent entites (eg. https://arxiv.org/abs/2510.01246). Until they publish something that turns out to be fraudulent, I think it's reasonable to consider Anthropic's paid people a very relevant, and in fact excellent part of interpretability discourse.

Dismissing all opinions where there is a perceived conflict of interest is a pleasant cognitive bias to have, but reality is often more nuanced than that.


I guess you don't really know how an LLM works then..?


10 years is a long time. 10 years ago the Transformer architecture didn't exist. I would call it moderately unlikely at best. At the very least, I would say it's likely that development will require an entirely different skillet 10 years from now.


There's at least one benchmark that attempts to measure this, but it has been running for a year plus so it's quite infrequently updated now.

https://fiction.live/stories/Fiction-liveBench-Mar-25-2025/o...


Ironically, the article points out that the original authors publisher actually put out two DMCA notices to google last year, apparently with no effect.

I guess DMCA takedowns are only for the big fish fighting the good fight against car pirates.


Why aren't they suing Google in court? Now is the best time to do this politically, the site doesn't mention what state they live in but I'd doubt that you wouldn't be able to get a state AG to listen to you if you reached out.

edit to add: Google has ignored all safe harbor protections, they would lose this protection and be held liable for all damages. This seems like a pretty solid win for the author here if they're telling the truth.


There are a lot of really solid reasons someone might choose not to sue one of the world’s most powerful corporations over their core business practices, even being completely in the right. Pretty similar to reasons someone might not call the cops on some mafiosos that are blocking your driveway while fencing a truckload of stolen electronics.


I disagree, state AGs are taking any opportunity to attack big tech for easy political points. If you want justice, now is the best time to try and get it.


Filling a complaint with the civil division of your AG so they can take legal action vs suing a giant powerful corporation and hoping the AG’s civil division steps in are very different things. I’ve also known numerous people who’ve filed multiple complaints with the MA AG with easily provable and well documented cases of repeated, ongoing wage theft as restaurant workers, at a time where service industry labor abuses were a popular hot button issue, and they didn’t even get a callback. Not exactly what you’d call a slam dunk. Not having support from the AG would be one of the very solid reasons I was referring to.


then DMCA the entirety of google and alphabet, and tender class action for direct and contributory violation, with the option to back it off to the literary work in question when goog takes a seat at the table and takes it seriously


That costs money


so does loosing a business.


Losing*


both spellings, while having differing implications, tend to be true.


Simon & Schuster is a small fish?


Eventually almost every other regulation turns out to be one that benefits big players and doesn’t help smaller ones.


Usually it’s easy enough to see who’s pushing for the regulation to figure out who will benefit from it.

Eg https://www.politico.com/live-updates/2026/06/16/congress/me...


The entire purpose of the DMCA is that it’s a bludgeon that the rich and powerful can use to beat the poor and powerless. It was never meant for individual copyright owners to use against giant infringing corporations.


DMCA is no longer valid as the courts ruled that stealing literally everything on the internet was OK and a valid business practice.


Once again, Russia turns out to be the reason we can't have nice things. War truly is a waste for everyone involved. Now that Russia is also helping North Korea to launch satellites (one so far), expect everything to get worse in the future.

I give it 2-10 years before one of the two threatens an imagined adversary with detonating a nuke in orbit, with the explicit intent of causing Kessler syndrome.


The article tiptoes around the who and the why but it's pretty clear.

The US has HARM (anti radiation missiles) but it would be very easy for Ukraine, or anyone, to create an inexpensive drone, which you would program with a frequency, a signal strength, and a general direction and send it to ride the radio signal straight to the jammer. Repeat daily as needed.

Anything on the ground transmitting high power on GPS freqs is up to no good for the global community.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: