You’re missing the point of the benefits of solutions like these, and the original set of tools like the Informatica of the kind. Those tools come with limitations and constrains, like a box of legos you can build a very powerful pipeline without having to wire up a lot of redundant code as you pass data frames between validation stages. Tools like Airflow/Spark etc are great for what they are, but they don’t come with guidelines or best practices when it comes to reusable code at scale, your team has to establish that early on.
You can open a pretty complicated large DAG in and right away you’ll understand the data flow and processing steps. If you were to do similar in code, it becomes a lot harder unless you comply to good modular design practices.
This is also why common game engine and 3d rendering tools come with a UI for flow driven scripting. It’s intuitive and much easier to organize.
This is indeed one of the nicest UI's I've seen, is this all custom made or are you using some open components? I wanted to get similar effect for something completely different I had in mind (like a flow based diagram), any advice greatly appreciated.
It actually IS quite hard to imagine solid engineers finding it hard to learn simple SQL. It's a lot harder to write "nasty" sql than "nasty" js, and if you're a front-end dev struggling with sql, then you probably should not be doing back end work. Yes not all databases are equal, thats why there are ORMs and ANSI standards.
From the list of bad practices you've "seen" people do in SQL, I would guess their comfort zone code is probably a hot mess as well.
> It actually IS quite hard to imagine solid engineers finding it hard to learn simple SQL.
I used to think this because I learned SQL right along with all my other coding. I've realized though it is a mindset shift to go from imperative to declarative, and to think mostly in set operations. That shift can be hard for otherwise good developers.
> It's a lot harder to write "nasty" sql than "nasty" js
That may be true (though, when it comes to dealing with more complex reporting functions and per-database-implementation differences/quirks outside of the realm of ANSI SQL it may not be).
However, it misses the point. Most engineers that struggle with SQL struggle with it because it's hidden from them partially (they're composing queries from snippets/query builders that come from other code, and never get to see the schema directly since it's hidden behind migrators and management interfaces) or completely (ORMs). Because the actual queries being run on actual schemas are less obvious, people do the wrong thing a lot.
That's not to say that abstractions on top of SQL are always bad--perhaps they're over/mis-used, but having worked on massive codebases where every dev's interaction with the database was "write a query in text or with a select().where().from()-type thin wrapper" and codebases that were 100% Django ORM, I can say with confidence that neither approach scales well absent big investments in correctness and RDBMS education.
Hi Rick, I've been following Memsql for a few years now, are there any plans to release "community" edition? Last time I checked about 1.5 years ago json support was very basic and EE pricing (dont remember exact #s) was rather high. Thanks
There already was a community edition[0]. And it was replaced by a "developer" edition that can no longer be used in production. It seems it didn't pan out as a marketing strategy and I don't think it's coming back.
EE features without support. I hate to bring up Mongo as example, but something similar...where support, additional software/plugins and cloud hosting are where the $ is made.
I did thorough testing of Memsql two years ago but went with Aurora instead. Would love to see how the product evolved since (Spark and Streaming integration was just being rolled out at the time), but something tells me pricing will be a deal breaker.
Thing is they are competing with PostgresSQL which you can extensively try for free before opting for a support.
"Free" being already hard to beat. The fact that you can't extensively test a solution is a real turn down for me (unless negotiating with commercials which is not nerds cup of tea).
I was asking because I am the creator of RediSQL[1] -- SQL steroids for Redis -- which is a less sophisticated product than MemSQL but still has its own use cases.
And maybe for parent was enough, or if not it would be very interesting to know what is missing.
Honestly, I believe that for small workload you can definitely use RediSQL in production, it will happily contain your cache or it will be a great SQL database.
However, I need a way to cut it between people just using the free product and people actually supporting the project, so provide as paying feature something that the big company will require it seemed to me the only way to go.
Unfortunately, I don't have the capital nor the bandwidth to go with fully open source product and selling just support, which I don't believe is anyway a good business model.
If you were in my shoes, you would do something different?
I haven't followed any links posted in this thread, but some things I see often are: free for non-commercial use, timed commercial use usually in the region of 30 days, or rates based on reads and writes. The last one seems like a winner from what you've described as your situation.
To be honest, I fail to see what I could use your product for so I'm out of the target audience.
Assuming nosql is for something very efficient or very scalable, I need some space to use it before I have to shell $$. There are many products where I have to pay before going on production.
If I were building a fast prototype I would not use a postgres box anymore but just a redis one.
If you need to cache data in a way more complex than just key->value you don't have too many alternatives at the moment.
If you want an easy and fast way to have an SQL engine in memory, again is not going to be simple.
If you need a separated database for every of your user there are no many alternatives that I am aware of.
It is definitely not a revolutionary product, but it has it's niche, any of the problems that I mentioned can be solved in a different way, but those different ways are quite complex.
Slava @ Rethink here. Just wanted to point out that only a tiny amount of credit for Horizon belongs to me, as most of it has been a product of tireless work of many people (hey Marc, Dalan, Josh, Daniel, Michael, Annie, Christina, Ryan, Mlucy, Chris, Marshall, and all the awesome beta contributors!)
This isn't false humility either -- I've been super busy with some other aspects of RethinkDB over the past few months, so while I did have a hand in designing and shipping Horizon, maybe less than 5% of it was my doing.
basically if you want fast ingress, keep shards small, once they get past ~5-10gb , ingress significantly slows down. Also this was on ES 1.5 , have not tested latest 2.0+ builds
if you want the fastest ingress, disable replica until your ingress is done, its faster to create replica at the end of ETL for that given index. Also, you want to disable auto allocation as well, this will disable shard movement during ingress, re-enable it afterwards.
on a 100 node cluster i had roughly 500GB on each node. this was not a single index, multiple indexes, with roughly 8 shards per index per node. Shard count is pretty important to get correct.
I did not manually control document routing (it was hard based on the type of data i was ingressing), so it was set to auto and during the load i observed hotspots in the cluster (you have to look at BULK thread/queue length), some nodes were getting burst of docs while others were idle, roughly 40-50% of the nodes in the cluster were under utilized, and maybe 5-10% had hot spots from time to time.
Also, depending what you use to push data in, (I used ES hadoop plugin) , you have to account for shard segment merges, which literally pause ingress for a brief moment and merge segments in a given shard. You have to set retry to -1 (infinite) and retry delay to something like a second or two, otherwise you will end up with dropped documents.