Hacker Newsnew | past | comments | ask | show | jobs | submitlogin
Intel Hyperscan is a high-performance multiple regex matching library (intel.github.io)
56 points by wslh on Sept 12, 2022 | hide | past | favorite | 20 comments


One of the Hyperscan developers commented on HN slightly over a year ago [1] that Intel had basically abandoned Hyperscan due to lack of sufficiently vocal customer interest. I hope that has changed, or is in the process of changing, because Hyperscan remains amazing.

[1] https://news.ycombinator.com/item?id=27421665


That post and the ensuing discussion are somewhat depressing. I am curious as to how the team's progress was measured; did the team's manager not know they needed to solicit user testimonials? Someone had to be responsible for the existence of the team and their productivity. What was the disconnect?


The official repo[1] is looking pretty dead.

Edit: This is a surprisingly heavy build. Over 15 minutes on my somewhat older i5. I suggest using make -j instead of cmake -build to speed it up. Decent number of compiler warnings too.

[1] https://github.com/intel/hyperscan


In sufficiently new versions of cmake you can use cmake --build -j as well. It will be translated to whatever argument is required for the underlying build system to use parallel compilation, e.g. -j for make and -m for msbuild.


It just seemed excessive for a regular expression library. For comparison the most recent version of PCRE built in almost exactly 26 seconds.

Sadly the hyperscan build crapped out for me in the linking phase, failing to find a bunch of coreX and avx optimized memset functions.


Not enough three letter agencies in the US to fund big brother at scale?

(One if the use cases of this library is scanning through bulk intercept data streams for a set of key words.)



In our firmware extraction suite, we started searching patterns with Yara, but it was so slow we switched to Hyperscan: https://github.com/onekey-sec/unblob/blob/main/unblob/finder...

We are quite happy with the huge performance boost, never looking back :) Yara was even segfaulted when we tried to extract an Ubuntu ISO image with it, never had problems like that with Hyperscan.

The only problem with Hyperscan is that it only supports Intel CPUs (it has some hand-crafted assembly), so it doesn't work on Apple M1 Macs, but there is a fork called VectorScan, which is working on ARM: https://github.com/VectorCamp/vectorscan

We implemented a couple of small DSL classes in Python to be able to define YARA-like rules, it might be useful for you too, check it out here: https://github.com/onekey-sec/unblob/blob/cdd7a46667ffdfdfae...


Second the support for vectorscan. And FWIW, I have posted issues in both hyperscan and vectorscan repos, and the only responses came from vectorscan (and quite quickly).


There's also Vectorscan[1], which is a fork of Hyperscan that runs on more platforms.

[1] https://github.com/VectorCamp/vectorscan


I'd be really curious to see how it optimizes for ARM. x86 is great for kitchen-sink optimizations like this, but with ARM you have a much smaller extension set to target. Hell, you'd probably be stuck with NEON unless you figured out a way to GPU-accelerate the process.


Esp. AMD, as Intel refused to accept AMD patches. So Hyperscan is effectively dead, the maintainer is also long gone.

VectorScan is its successor.


C# has also recently put a lot of work into speeding up regex. MS put up a post about it a while ago going into the specifics. I expected the vectorization. I did not expect source code generation.

https://devblogs.microsoft.com/dotnet/regular-expression-imp...


Seconding the comment, .NET 7 has seen massive improvements to Regex performance on all platforms supported by .NET. Char sequence matching will use respective SSE4/AVX2/AdvSimd instruction sets and for compiled regexes, efficient automata will be emitted. Those two combined produce state-of-the-art regex matching performance.


Related:

Hyperscan: High-performance multiple regex matching library from Intel - https://news.ycombinator.com/item?id=21873557 - Dec 2019 (53 comments)

Hyperscan: A Fast Multi-Pattern Regex Matcher for Modern CPUs - https://news.ycombinator.com/item?id=19270199 - Feb 2019 (36 comments)

Multiple regex performance shootout: RE2 vs. Intel's Hyperscan - https://news.ycombinator.com/item?id=14608663 - June 2017 (55 comments)

Hyperscan, a high-performance multiple regex matching library - https://news.ycombinator.com/item?id=10420295 - Oct 2015 (19 comments)


PSA that the python bindings for hyperscan need maintainers: https://github.com/darvid/python-hyperscan/issues/44


Capturing is not supported. That eliminates many normal regex usecases, where you want to be able to capture & interpret the data.


Capturing can be supported by using another regex library in conjunction. Use Vectorscan/Hyperscan to eliminate most of the search space, then run a (slower) more featureful regex matcher where the former tells you is a match to extract information.

After all, Hyperscan's original objective was to be able to find where, in a large input, any of many patterns may match. It doesn't really make sense to use it as a general purpose regex library, because it wasn't built for that.


At scale though?

And even if that were the case you could grab your sentences from the 1TB file in front of you with hyperscan, and then do a capture on the results.


Somewhat related: I built a library to speed up matching many regexes with mostly mismatches by adding non-regex pre matchers. https://github.com/Quantco/multiregex




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: