Fast deduplication data transmission scheme on a big data real-time platform

Sheng Tzong Cheng, Jian Ting Chen, Yin Chun Chen

Research output: Chapter in Book/Report/Conference proceedingConference contribution

Abstract

In this information era, it is difficult to exploit and compute high-amount data efficiently. Today, it is inadequate to use MapReduce to handle more data in less time let alone real time. Hence, In-memory Computing (IMC) was introduced to solve the problem of Hadoop MapReduce. IMC, as its literal meaning, exploits computing in memory to tackle the cost problem which Hadoop undue access data to disk caused and can be distributed to perform iterative operations. However, IMC distributed computing still cannot get rid of a bottleneck, that is, network bandwidth. It restricts the speed of receiving the information from the source and dispersing information to each node. According to observation, some data from sensor devices might be duplicate due to time or space dependence. Therefore, deduplication technology would be a good solution. The technique for eliminating duplicated data is capable of improving data utilization. This study presents a distributed real-time IMC platform - "Spark Streaming" optimization. It uses deduplication technology to eliminate the possible duplicate blocks from source. It is expected to reduce redundant data transmission and improve the throughput of Spark Streaming.

Original languageEnglish
Title of host publicationBMSD 2017 - Proceedings of the 7th International Symposium on Business Modeling and Software Design
EditorsBoris Shishkov
PublisherSciTePress
Pages155-166
Number of pages12
ISBN (Electronic)9789897582387
DOIs
Publication statusPublished - 2017
Event7th International Symposium on Business Modeling and Software Design, BMSD 2017 - Barcelona, Spain
Duration: 2017 Jul 32017 Jul 5

Publication series

NameBMSD 2017 - Proceedings of the 7th International Symposium on Business Modeling and Software Design

Other

Other7th International Symposium on Business Modeling and Software Design, BMSD 2017
Country/TerritorySpain
CityBarcelona
Period17-07-0317-07-05

All Science Journal Classification (ASJC) codes

  • Modelling and Simulation
  • Software

Fingerprint

Dive into the research topics of 'Fast deduplication data transmission scheme on a big data real-time platform'. Together they form a unique fingerprint.

Cite this