You are right, but what I meant is that delta between Spark and Hadoop is RDD (RDD - Resilient Distributed Dataset) which is data structure first of all. Thus Spark and Hadoop have the same data processing model implemented over different data representation models (and hence the performance gain).