Not sure if this is a dumb question, but if your cache is behaving so badly (40% 500 error?), shouldn't you just get rid of the cache? Maybe allocate some more resources to the DB, but the DB already caches frequently accessed data and you save yourself the roundtrip of check-cache-then-db.
Original author here. I think i wrote the sentence in the post confusing.
The error rates of (40% 500 error) were not normal.
A normal hit rate was at ~97%. But because the redis server itself was overloaded by KEYS request, BGSAVE / forking, etc. this instance was not able to answer the cache request.
Due to this we had a high error rate.
But in general i agree. A normal behaviour of a cache with such a bad rate should be avoided.
I hope this make more sense now?
The conventional wisdom is to move as much load as possible from DB to a cache. The reason is quite simple. Scaling a database is hard and expensive. Scaling a cache is very easy in comparison.