Hacker Newsnew | past | comments | ask | show | jobs | submit | doctorpangloss's commentslogin

> For example, there is a 100% test coverage requirement in both server and client, combined with AI-driven review rules that say all tests must be non-vacuous,

look, your application works, right? so it doesn't really matter what you or i think, and this is why AI matters. but this, your "100% test coverage" - that is pure slop. just 20 years ago, all the most popular software shipped with NO tests. are you getting it?


> just 20 years ago, all the most popular software shipped with NO tests

not sure what you're point is here. It sounds similar to "we use to use blood letting and leeches and doctors didn't clean their hands and everything was fine so what are you getting at?"

Good tests have real benefits. The fact that people shipped without them in the past in no way suggests they aren't needed or have no point.


Also, 20 years ago SW was tested by QA department and approved before shipping. Don't want to go back to that, but there were testing, just differently.

IMHO getting rid of proper QA done by teams of QA specialists is the main reason for the current software quality crisis (and that already started 15 years ago or so). We should go back to QA teams and proper QA procedures! Automated tests are no replacement, especially when they are set up by the same people designing and building the product.

Well. First they came for the testers. Then they came for the developers.

> just 20 years ago, all the most popular software shipped with NO tests

Are you getting older? A lot of people anchor their intuition of time and history to a certain year. There are probably still lots of people who think the 1990s is not that long ago even though it’s now over a quarter century since it ended. Maybe you mentally default to 2012 or so, when it might be true that most popular software shipped without automated tests (although manual QA was a lot more extensive in 1992).

But 20 years ago is now 2006, and unit tests were well established as a best practice. Perl had extensive automated tests in the late 1990s that everyone who ever compiled Perl would have noticed, since they were run by default and produced obvious output. Kent Beck’s “Test Driven Development: By Example” was released in 2002, and popularized both the name and practice.


But 20 years ago is now 2006, and unit tests were well established as a best practice.

I agree with the spirit of what you wrote, but my recollection of the timeline is different. The first decade of the 2000s was peak Crazy Agile Advocacy, but IIRC it wasn’t until the 2010s that unit testing really became almost universal practice. Much before that and it was still tangled up with XP, TDD and lots of other things that certainly weren’t universally accepted as good practices (notwithstanding the strident advocacy of a certain group of consultants/authors/speakers/bloggers and their fans).

I remember, back in the mid-2000s, when we had some consultants brought in to talk about different aspects of quality and testing. There were several working groups, each led by one of those external consultants, and one of them was about unit testing. This was in a relatively large software development organisation for the time, a few thousand people, and while some parts of the organisation had some form of automated testing operating by then, it definitely was not the case that the well-known products produced by the organisation all had a unit test suite. Other practices we’d consider routine today, such as peer code reviews, were also in their infancy during that period: some were doing them, many were not, and generally we had much less experience of how to do them effectively than we have today.

As an industry, I don’t think we really matured in how even the most ardent fans of unit testing were writing test suites until the 2010s either. In the 2000s, we still had lots of people mocking the entire universe and then writing unit tests that were 99% testing those mocks because of 100% test coverage requirements, and similar dogmatic nonsense.

By the 2020s, I think there was much more awareness of that automated testing is generally a good idea, but there are different kinds/levels of automated testing and finding a mix that suits each project’s specific needs is important. One of the great benefits from the more recent AI tools, particularly the agentic ones over the past year or so, has been that it has clearly demonstrated both the value of a good automated test strategy and how much of a waste of time vacuous tests are.


What kind of software are you talking about? Test harnesses were commonly used in 2006. JUnit was created in 1997.

> just 20 years ago, all the most popular software shipped with NO tests. are you getting it?

Not really?

About 20 years ago, I was working on Firefox and we had millions of tests on CI. I was working on a host of other open source apps and they all had tests (most of them had no CI, of course).


And 80 years ago cars didn't have seatbelts. Your point?

Careless drivers got weeded out.

Cars today are safer than ever. Drivers (in my memory anyway) have never been worse.


But surely then you agree they are having an effect. Thus my analogy serves its point - test coverage isn't to be dismissed out of hand.

If only the careless drivers were those crippled or killed by their poor driving, it'd be a self-resolving issue as you imply, and nobody should care.

However, what actually happens is that careless drivers often cripple or kill innocent bystanders in other vehicles as a result of their poor driving. That's why seatbelt laws and improved vehicle safety features are a good thing.


>> but this, your "100% test coverage" - that is pure slop.

Not really, but I can see why some people think that.

We treat 100% test coverage as "required, but by itself not sufficient". It doesn't give us false confidence that everything will be perfect or anything like that. But it provides us with the discipline to make sure no corners are cut, and the bugs that are fixed don't come back.

One refreshing aspect was that during PR reviews we stopped debating whether something needed test coverage. Instead we focused on what was being tested and how.


> One refreshing aspect was that during PR reviews we stopped debating whether something needed test coverage.

I’d be curious to know what percentage of the time spent implementing tests would have otherwise gone to discussions about whether to implement them or not. ;)


the requests went into a pipeline that turns them into de-identified, but salient, training data, yeah. everywhere except maybe bedrock.

It's complicated.

SMBs can find an audience that would otherwise be impossible to reach. The thing that is advertised is so diverse it's hard to generalize, however the main thing I've noticed is SMBs live and die by the quality of their ad creatives.

For big businesses which make up the other 50% of Meta ads revenue: They are very effective in circumstances where your main offering is operational and the long term value of a customer is very large. For example it would be a very good place to advertise health insurance and TV subscriptions (which is a significant amount of the big business spend). Here I am more certain. If you didn't invent your product, and it is a beneficiary of the zeitgeist promoted by social media itself (beauty-focused, soft core brain rot), and your thing is basically fungible, you will thrive.


I don't know. Apple marketing has been really effective at this health thing. Whereas, studies of continuous monitoring of eg heart with even better sensors and higher risk populations have shown conclusively no improvement in health outcomes.

Yeah, I only see anecdotal evidence of Apple Watch helping people: individual heavily-promoted stories.

Still, on a personal level, having my health data over time is something I find interesting to collect.

Things like my sleep, gym heartrate (I mostly pay attention to zones rather than specific numbers), sleep heartrate, occasional ECG reading, steps, with added smart scale and EMR data… It does show interesting trends over time.


I would have guessed the opposite.

Could you share a couple of these studies you are citing? Curious to read more.


> If they opted out of training, then we definitely did not train on them.

are opted-out-of-training chats ever paraphrased, and thus "de-identified" (in openai's own words)? what this means is, it would be hard to prove that a "synthetic" (but actually paraphrased) training instance came from a particular chat. except if the chat was about some esoteric math proof, of course.


> It would be extraordinarily easy to simply say, this model was not trained on your work, if that were the case.

well, it is trained on their work. all user inputs are paraphrased for training. at openai, at anthropic, at google, and now with all the bedrock models, and at openrouter providers, even if they say zero data retention.


> While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models

it just means that they've been working with paraphrased user data everywhere, which any smart person can figure out is how anthropic and openai train on so called non-retained data.


the reality is that those stats were flawed in numerous ways. this is expected, they were never subject to peer review.

for example: ok cupid and tinder, during the times that christian rudder was blogging about them, never reported on the actual ratio of active male to active female users. this made it impossible to do any comparisons over time, it's like reporting only numerators in the absence of denominators.

a second example: when tinder's 2018 interracial dating stats were compared to okcupid's 2011 stats, there was no change, despite real preferences having shown that something like 40 percentage points more people were comfortable with non-same-race dating over that decade. rudder kind of failed the obvious interpretation, that the apps and/or data collection were broken, and instead talked about some conspiracy of the surveys being wrong.

i think there was a lot of other stuff broken with it.


As a young adult that data analysis blog post seemed like cool shit, and today I'm horrified because it seemed like cool shit. They make for "cute" PR soundbites, and back then the data analysts likely didn't know better, but they are terribly biased and reductive by virtue of the structure of the platform, and can't be used to support anything essential or fundamental about human romance.

The most valuable data is the per user purchase histories from before the time those were hidden by default.

Why is that "valuable"?

they basically map out the paying audience for your game on Steam

So we know how many hentai games you've bought

i don't know. the whole operation was kind of obnoxious. valve could provide all this information for free in an afternoon, but chooses not to. of course, by information, i mean: historical, per-user purchasing histories, which are 1,000,000% the most valuable piece of data. but since the guy who ran steamdb doesn't make games, he didn't comprehend what it was that was valuable, and instead, focused on CSGO "leaks."

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: