This is a cute example, but it misses the mark. The efficient way to do fast nearest neighbor search is with a search tree (e.g., KDTree or BallTree), which brings down query time from linear to logarithmic in the number of items.
Agreed, but since those two concerns are separate (algorithmic improvement vs. taking advantage of parallel hardware) I'm not sure I'd categorize this as "missing the mark" so much as that's a further improvement that could be made. For a blog post that looks to be an attempt to tout methods to easily exploit data parallelism, I think focusing on algorithmic improvement would be counter productive.
I tend to agree, but there is a tendency to throw hardware at things where some algorithmic improvements would work much better. See for example this blog post
I was a bit disappointed that there wasn't much detail given about the problem being attacked, and no information about the results other than some timings. Sure, Julia's nice, but what are we looking at here?
Also, there's talk about how slick this is when using IJulia notebooks. It would be cool to provide a link to an actual notebook.
It would be great to see the difference with other languages. So why Julia and not R, or Matlab, or Python? Is it more elegant, more concise, does it have more libraries, can it be run in parallel better? That would be great to know!
Julia is a lot (lot) faster than Python, Matlab or R. This doesn't matter so much if you're gluing library calls together (if using well-known ML algorithms that use native BLAS/LAPACK etc) but for custom stuff Python et al are just too slow. Julia is comparable to C++ in speed (not as fast, but within an order of magnitude in my experience) and it's MUCH more fun to write.
There are definitely not more libraries though, it's still a young project.
The problem is: plain Julia is faster than plain CPython. But Numba has solved this problem for me, by compiling my hot loop in-place. Other people report similar success with PyPy, Cython, or Pyston.
Furthermore, libraries like scikit-learn, pandas, matplotlib, or scipy are incredibly powerful, and usually implement in fast languages. Python is only there to glue them together.
For my applications, I just don't see any compelling reason why I should use Julia over Python. In practice, Julia is (for the above reasons) not faster in practice, and the libraries are a lot less mature. I try Julia every few months though, and there is progress. Maybe in a few years.
Between ccall, PyCall and RCall you can access a huge range of existing code from inside Julia. So the libraries issue is moot.
Having used both, I much prefer Julia to python+cython. Three examples: multiprocessing is much less restrictive, there's no edit-compile loop so development is quicker, and no awkward pyx/pxd system. If you have to go deep with a cython project, you basically end up writing C, and that's slow - not cython's fault, it's a great tool, just a limitation of that platform.
A quick look at the numba docs suggests that it doesn't parallelise any better than regular python - so there's one of my problems.
Numba also seems to be restricted in the functionality it offers, e.g. currently looks like no user defined types so good for hot loops but not your whole codebase. If you touch the python C API it suggests it can't do full JIT compilation, i.e. sounds like this would happen with any C-extension library code outside the range of supported numpy features. No strings?...
If you're really happy with your current tools, Julia's probably not for you right now, and that's totally fine. Julia's users tend to have the opposite of the "established libraries + glue" use case; something more like "custom data structures + unvectorisable numerics". It's better for the people who are writing the next scikit-learn than for those who are using it.
The main advantage of Julia's JIT is that the performance of a given piece of code is easy to predict and debug – this is a classic problem with tracing JITs.
Julia has a lot of room for improvements. For example with better type inference (programs get slooow when the type cannot be inferred, they are working on it) and stack-allocation of arrays.
Currently, all arrays require a malloc. Once the JIT-compiler is sufficiently smart, Julia could use stack-allocation of temporary arrays (that cannot escape the current context). Algorithms making use of temporary arrays will become faster by possibly an order of magnitude.
Julia might become competitive with C / Fortran for general purpose computing. And that is a level of speed all your other examples cannot hope to reach.
sort of: Julia is JITted, so there can be some overhead, and if you're not careful about how you program, the overhead can be high (and it's not hard to be careful), but the cost of this overhead diminishes as you run over increasing amounts of data. Also some things like Julia's text IO are super slow because of the architecture of how functions like print() work (I think this is a work in progress) so logging may incur a cost.
Don't get me wrong, I love Julia and actually get paid to code in Julia, and it's a real joy of a programming language to work in (I'd put it close to ruby in terms of programmer satisfaction)... Just think that overblowing speed claims is counterproductive.
> Don't get me wrong, I love Julia and actually get paid to code in Julia, and it's a real joy of a programming language to work in (I'd put it close to ruby in terms of programmer satisfaction)... Just think that overblowing speed claims is counterproductive.
Whoa! Who is using Julia in production? I'm a data scientist in a large corp that has successfully converted folks to Python, and would love to use Julia for my custom stuff!
I wouldn't call it 'production'. I'm prototyping experimental number systems for high-performance computing and my supervisor has been doing it in mathematica to date, which is so glacial it's unacceptable. Took hours to do a calculation that julia can do in seconds.
It's very elegant. The type system and multiple dispatch are a little bit tricky to wrap your head around if you're coming from an OO language - and especially if you're coming from a duck-typed, monkey patching, OO language. But once you get it, it's actually simpler, and makes way more sense, and there are real performance advantages.
I'd say it's more readable. Here's why I like to read Julia code more than Python code (and Python code is normally already pretty readable):
* multiple dispatch in Julia allows to write shorter functions for each specific type and type parameters, while in Python you have to check them all in function's body. E.g. `numpy.dot(a,b)` needs to check type and number of dimensions for `a` and `b` in the same place, while in Julia `` is overloaded for each pair of arguments.
composite types in Julia normally include all fields with their types during declaration. In Python fields may be added wherever in class definition, types are not specified at all
* Julia has much broader set of overloadable operators (e.g. dot-operators like `.*`). Together with multiple dispatch it makes creating custom data types (e.g. custom array types) much easier.
> does it have more libraries
Definitely not, but it has very good relations with other languages. Calling a C function boils down to one line of code, calling Python or Java - couple of lines, never tried to call R or Matlab, but it seems to be fun too. In general, I've got much more pleasant experience than with any other pair of languages.
> can it be run in parallel better
Julia support (1) concurrency via tasks/coroutines, (2) shared-memory parallelism via threads (v0.5 only) and (3) isolated multiprocessing on local and remote machines. Most other scientific languages support either only (3) or (1) and (3) (Python 3+ version).
There's also a couple of unique features in Julia. My favorite is metaprogramming support. For example, currently I'm working on a library for symbolic differentiation from source code - something that would be very hard to achieve without direct access to AST. Also, macros allow to reduce boilerplate code a lot and make it easy and straightforward to create DSLs.