Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

> That's why I'm a skeptic of a number of modern development practices that deliberately increase the complexity of software

Maybe I'm nitpicking, but the article points out the number #1 predictor of software bugs is not the complexity of software but of the organization itself. A single person can make an hugely complex piece of software, and a relatively large team can make a conceptually simple system.

As for software complexity itself, there's an interesting research result that the thing that matters the most is line count. Not cyclomatic complexity, not the type system, not modularity or test coverage, not the programming language -- looking at the line count alone trumped all the other metrics in predicting flaws. (I can't look for this paper now, but I'm sure with a bit of googling anyone can).



A single person can make an hugely complex piece of software, and a relatively large team can make a conceptually simple system.

In practice I think you tend to hit Conway’s law -- organizations build software that mirrors their own organizational structure. So it’s hard for large teams to make simple designs.

I’m very skeptical about that line count metric; in my observation bugs tend to sit between modules due to bad interface design. But I could certainly believe that bad modularisation is correlated with line count (in the form of excessive boilerplate).


You can be skeptical, but this has been tested empirically. Defect density per lines is relatively constant (this is a well-known empirical result, which you can also google), which means more lines predict more bugs. When researchers contrasted line count vs bad modularization, line count was the better predictor (precision & recall).

Maybe it's anti-intuitive, but it's what reality shows :)


Maybe it's something as banal as there being a high enough chance of mistakes per LOC that this factor ends up dominating the other factors.


Maybe. The key takeaway of this for me is: if you want to introduce a bug-predicting metric, you must do better than simple LOC count. If you can't do better, then ditch your metric :)


>if you want to introduce a bug-predicting metric, you must do better than simple LOC count

Agreed. As I pointed out in another post[0], a good candidate for this metric is the amount of unnecessary code, which is often a proxy for unnecessarily large/complex teams.

[0] https://news.ycombinator.com/item?id=21824110


There's also just reading/remembering costs per LOC that make it harder to track specifics about APIs and your own code.

Smaller code fits in your head better, and stays more predictable. You don't have as many weird if-conditions to remember.


You don't have as many weird if-conditions to remember.

More LOC doesn’t necessarily mean more if-conditions, though! For code written in different styles.

It’s a strong statement that LOC and LOC alone is the best known predictor of bug rate, but that seems to be what several people in this thread are saying.

For example, I think most people (though not all) would agree that very dense code full of tricky one-liners (think Perl Golf) is more likely to be buggy and hard to maintain. But if so, “number of control structures used” or some such ought to be a better metric than plain LOC (I assume we’re using “LOC” literally here, not as a shorthand for something else).

Maybe there’s some second-order effect going on? Like very dense code discouraging modifications, so it gets less maintenance, and therefore accrues fewer bugs over time?

This is an interesting topic! Any research links appreciated.


You can be skeptical, but this has been tested empirically.

Yeah, I wasn’t claiming to be right, more noting my reaction! I’ll look up the research -- any links appreciated.


> in my observation bugs tend to sit between modules due to bad interface design

Yep, and when you start ripping the monolith apart into separate version controlled projects and deployment pipelines without addressing the interface issues you've significantly increased the complexity of your work products.


Keep in mind that size doesn't necessarily equal complex. Where I have seen Conways law most clearly is with dysfunctional teams regardless of size.

I was on a team once where one of the senior people was such a jerk no one wanted to work with him. This led to him and the team carving out a piece of the system that he alone worked on and interfaced with the rest of the system through a single queue. This was certainly not the best design and added all sorts of unnecessary complexity.

Another team I was on had a person who was a good programmer and tended to blow off design meetings. The organization never rarely reprimanded this person. In turn it led to various APIs being built that were close, but always not quite right.

Big company examples abound. Contrast and Apple keynote to a Google one. Sometimes I wonder if the people presenting at the Google one even work at the same company.


That makes sense. Every line is a potential bug.

But it's not the only factor, and, quite frequently, it's a matter of correlation, as opposed to causation.

That kind of thing can be very tricky to determine.

When I write software, my first stab at a function tends to be a fairly linear, high-LoC solution, which I then refactor in stages; reducing LoC each time, and ensuring that the quality level remains consistent, or improves.

As far as quality goes, my first, naive stab, was just fine, and I have actually introduced bugs during my refactoring reduction.


Note that I'm not arguing about underlying reasons, just saying what the empirical results show. That's reality.

So if you try to predict software bugs using modularization (or lack thereof) or "if-then-else" branches, or whatever complexity metric you can think of, you'll get one result. If you try to predict them using simple line count, you'll get another result. The second one will have better precision & recall. So no metric so far has been shown to be better than simply counting lines. That's not an obvious result, but it's the truth.

Sadly, you'll have to believe me because I cannot find the studies right now.


Any results/findings for development speed or new hire on-boarding?


> Every line is a potential bug.

But there are many code lines that don't contain bugs. If only one could somehow make software only from those...

I'd take a slightly more mathematical approach to the code lines predictor: zero lines of code contain zero bugs, an infinite amount of lines of code contains an infinite number of bugs. It follows immediately that yeah, LOC is massively important and no, we are not interested in that, what we do want to know is how everything else is influencing the derivative.

(and on that digression about longish linear solutions: I completely stopped feeling bad about writing long, linear functions for long, linear problems. I've come to greatly prefer those over the indirections of forced subdivision. If there's a reason to subdivide other than "long is bad", great, go for it, but never subdivide just for having smaller parts. Use assign-once, nested scope etc to make the length more palatable)


Pretty sure when the tests are green you're just meant to reach for the next sticky note ;)


Do you know if the paper controlled for all the variables in tandem. For example, I wouldn't think line count is necessarily mutually exclusive to modularization. That is to say, I'd expect well modularized code to reduce line count generally.


I found a video about the paper: the presenter claims the author of the paper (not the same person) did a double-barrelled correlation and that no metric had better predicting value than "simply doing wc -l on the source code".

See minute 39:30: https://vimeo.com/9270320


Thanks for finding that.


You're welcome. I've watched the whole video and it's pretty interesting. Not being Canadian or Australian I missed most of the jokes though :P


I think it did, that's the thing. Unfortunately since I cannot find the paper now, I can't tell you :(




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: