One of the Hyperscan developers commented on HN slightly over a year ago [1] that Intel had basically abandoned Hyperscan due to lack of sufficiently vocal customer interest. I hope that has changed, or is in the process of changing, because Hyperscan remains amazing.
That post and the ensuing discussion are somewhat depressing. I am curious as to how the team's progress was measured; did the team's manager not know they needed to solicit user testimonials? Someone had to be responsible for the existence of the team and their productivity. What was the disconnect?
Edit: This is a surprisingly heavy build. Over 15 minutes on my somewhat older i5. I suggest using make -j instead of cmake -build to speed it up. Decent number of compiler warnings too.
In sufficiently new versions of cmake you can use cmake --build -j as well. It will be translated to whatever argument is required for the underlying build system to use parallel compilation, e.g. -j for make and -m for msbuild.
We are quite happy with the huge performance boost, never looking back :) Yara was even segfaulted when we tried to extract an Ubuntu ISO image with it, never had problems like that with Hyperscan.
The only problem with Hyperscan is that it only supports Intel CPUs (it has some hand-crafted assembly), so it doesn't work on Apple M1 Macs, but there is a fork called VectorScan, which is working on ARM: https://github.com/VectorCamp/vectorscan
Second the support for vectorscan. And FWIW, I have posted issues in both hyperscan and vectorscan repos, and the only responses came from vectorscan (and quite quickly).
I'd be really curious to see how it optimizes for ARM. x86 is great for kitchen-sink optimizations like this, but with ARM you have a much smaller extension set to target. Hell, you'd probably be stuck with NEON unless you figured out a way to GPU-accelerate the process.
C# has also recently put a lot of work into speeding up regex. MS put up a post about it a while ago going into the specifics. I expected the vectorization. I did not expect source code generation.
Seconding the comment, .NET 7 has seen massive improvements to Regex performance on all platforms supported by .NET. Char sequence matching will use respective SSE4/AVX2/AdvSimd instruction sets and for compiled regexes, efficient automata will be emitted. Those two combined produce state-of-the-art regex matching performance.
Capturing can be supported by using another regex library in conjunction. Use Vectorscan/Hyperscan to eliminate most of the search space, then run a (slower) more featureful regex matcher where the former tells you is a match to extract information.
After all, Hyperscan's original objective was to be able to find where, in a large input, any of many patterns may match. It doesn't really make sense to use it as a general purpose regex library, because it wasn't built for that.
Somewhat related: I built a library to speed up matching many regexes with mostly mismatches by adding non-regex pre matchers. https://github.com/Quantco/multiregex
[1] https://news.ycombinator.com/item?id=27421665